LabsAmazon
Amazon
Track Amazon model scores across public AI benchmarks including GPQA Diamond, SWE-bench Verified, MMMU-Pro, BFCL, and MCP Atlas. Each result is shown one benchmark at a time, with source links and evaluation dates — no blended score or composite ranking. 2 models tracked, spanning Amazon Nova (Nova 2).
Models 2
Amazon Nova 2 Lite
Amazon Nova (Nova 2) · amazon nova 2 lite
2025-12-02
7 results
Amazon Nova 2 Pro
Amazon Nova (Nova 2) · amazon nova 2 pro
2025-12-02
6 results
Progress by benchmark
Show progress on
Single benchmark only
This view shows GPQA Diamond (accuracy) only. Other benchmarks use different metrics and are not directly comparable.
Progress matrix
Scores are not normalised across benchmarks. Each column uses its own metric. Compare columns independently.