Thanks for your excellent work!
In your paper, Qwen3-8B was used as the reference model. However, its test pass rate is quite low on the Terminal-Bench-2 benchmark. I was wondering if you have also evaluated the Qwen3.5 series as baselines to assess the performance improvement brought by SFT on the Terminal-Bench-2 benchmark.
Looking forward to your response!
Thanks for your excellent work!
In your paper, Qwen3-8B was used as the reference model. However, its test pass rate is quite low on the Terminal-Bench-2 benchmark. I was wondering if you have also evaluated the Qwen3.5 series as baselines to assess the performance improvement brought by SFT on the Terminal-Bench-2 benchmark.
Looking forward to your response!