A 7-point Terminal-Bench jump from 1,200 examples
4 days ago
One paper today. FACET, from Shanghai AI Lab, U S T C, and Fudan. It answers a question every coding agent team has right now. Where does trustworthy training data for terminal agents come from, when writing tasks by hand does not scale. Here is the first pass.
Ask
Ask about this presentation
Answers are generated from this presentation.