LabsSakana AI
Sakana AI
Track Sakana AI model scores across public AI benchmarks including SWE-bench Pro, Terminal-Bench 2.1, HLE, GPQA Diamond, and CharXiv. Each result is shown one benchmark at a time, with source links and evaluation dates — no blended score or composite ranking. 2 models tracked, spanning Fugu.
Models 2
Progress by benchmark
Show progress on
Single benchmark only
This view shows SWE-bench Pro (% resolved) only. Other benchmarks use different metrics and are not directly comparable.
Progress matrix
Scores are not normalised across benchmarks. Each column uses its own metric. Compare columns independently.