LabsStepFun
StepFun
Track StepFun model scores across public AI benchmarks including SWE-bench Verified, SWE-bench Pro, Terminal-Bench 2.1, AAII, and GDPval-AA. Each result is shown one benchmark at a time, with source links and evaluation dates — no blended score or composite ranking. 1 models tracked, spanning StepFun Step series.
Models 1
Progress by benchmark
Show progress on
Single benchmark only
This view shows SWE-bench Verified (% resolved) only. Other benchmarks use different metrics and are not directly comparable.
Progress matrix
Scores are not normalised across benchmarks. Each column uses its own metric. Compare columns independently.