Open SOTA registryRegistry snapshot 2026-04-27

Find the best AI model
for the job.

Independent benchmark evidence across models, tasks, and modalities. Every published score is dated, source-tiered, and linked back to where it came from.

A leaderboard row is not a fact until it can be inspected.
Try
9,102Results
163Models tracked
371Benchmarks indexed
121Tasks
9Capability areas
70Published SOTA
§ 01 · Inspectable frontier

What can the evidence support right now?

One protocol-specific leader per benchmark. A model is printed only when the row is verified and its source URL is inspectable.

General document OCROCRBenchqwen3-5-397b-a17bscore931.0✓ Inspect source
Multilingual document OCROCRBench v2 · Englishovis2-5-9boverall-en-private63.4✓ Inspect source
Document parsingOmniDocBenchGLM-OCRcomposite94.62✓ Inspect source
PDF conversion qualityolmOCR-Benchinfinity-parser2-propass-rate87.6%✓ Inspect source
Terminal coding agentsTerminal-Bench 2Codex / GPT-5.5accuracy82.0%✓ Inspect source
Software engineering agentsSWE-bench VerifiedClaude Opus 4.7resolve-rate87.6%✓ Inspect source

Leaders are computed only within the named metric and protocol. Incomplete or malformed evidence is withheld rather than silently promoted.

§ 02 · Choose by task

Start from the job, not the model.

§ 03 · Evidence

A leaderboard is only as good as the trail behind it.

Measured results, operational claims, and source quality stay separate so a clean-looking table cannot hide a weak comparison.

Every public number carries its audit trail.

CodeSOTA preserves metric direction, the exact benchmark variant, source type, and snapshot used. A vendor claim can be useful, but it is never presented as an independent reproduction.

  • Verified — inspectable evidence meets the publication floor
  • Vendor-reported — primary claim, not independently reproduced
  • Reproduced — an independent run and protocol are available
  • Withheld — the available row does not meet the evidence floor
Read the methodology →
§ 04 · Decision quality

A top score is the start of a decision, not the end.

Benchmark health tells you whether a table still separates models. The decision workspace then adds hardware, openness, language, and evidence constraints.

Matched-snapshot signal
Separating

ARC-AGI-1

31.4-point spread in the current matched frontier snapshot.

Protocol-sensitive

SWE-bench Verified

A strong signal only when harness and scaffold details travel with the score.

Low separation

ARC-Challenge

0.8-point spread across the matched three-model snapshot.

Saturated snapshot

GSM8K

No spread in the matched snapshot; use harder successors for frontier choices.

These are observations from one matched CodeSOTA snapshot, not permanent benchmark properties. Inspect the research question →

§ 05 · Original research

The registry remembers how a result became a claim.

This is the useful part of the older CodeSOTA homepage: public experiments, negative results, and lineage—not just another table of scores.

Public research graph

Question → experiment → evidence → claim → finding

Follow a decision back through its hypotheses, protocol, evidence, corrections, and negative results. The graph keeps research memory inspectable instead of flattening it into a marketing number.

Explore the research graph →
Questions
7
Experiments
14
Claims
18
Findings
6
§ 06 · API

The same registry, callable by software.

Agents should not scrape leaderboards. Call a stable, source-aware endpoint and retain the snapshot id so the number in a report can be reproduced later.

  • Public, CORS-open JSON with task-first rankings
  • Benchmark, metric direction, result date, and snapshot id
  • Short task aliases for OCR, code, ASR, TTS, and more
Read the API docs →
GET /api/sotacurl "https://codesota.com/api/sota/ocr?tier=sota"

{
  "task": "ocr",
  "tier": "sota",
  "pick": {
    "model_name": "ovis2-5-9b",
    "score": 63.4,
    "score_metric": "overall-en-private",
    "higher_is_better": true
  },
  "snapshot_id": "2026-04-27"
}
§ 07 · Registry loop

Maintained in public—and open to correction.

Full activity log →
  • 9,102 results tracked across 163 models

  • + 14 public experiments retain their protocol and outcome

  • ± Superseded claims remain visible in the provenance trail