Open SOTA registryRegistry snapshot 2026-04-27

Find the best
model for the job.

Independent benchmark evidence across models, tasks, and modalities. Every published score is dated, source-tiered, and linked back to where it came from.

Try
9,102Results
163Models tracked
371Benchmarks indexed
121Tasks
9Capability areas
70Published SOTA
Observed use · rolling 12 monthsNot a projection.Production analytics · through 4 Sep 2026
154,945Visitors
394,299Page views
2.54Views per visitor
Full traction record →
§ 01 · Inspectable frontier

What can the evidence support right now?

One protocol-specific leader per benchmark. A model is printed only when the row is verified and its source URL is inspectable.

General document OCROCRBenchqwen3-5-397b-a17bscore931.0✓ Inspect source
Multilingual document OCROCRBench v2 · Englishovis2-5-9boverall-en-private63.4✓ Inspect source
Document parsingOmniDocBenchGLM-OCRcomposite94.62✓ Inspect source
PDF conversion qualityolmOCR-Benchinfinity-parser2-propass-rate87.6%✓ Inspect source
Terminal coding agentsTerminal-Bench 2Codex / GPT-5.5accuracy82.0%✓ Inspect source
Software engineering agentsSWE-bench VerifiedClaude Opus 4.7resolve-rate87.6%✓ Inspect source

Leaders are computed only within the named metric and protocol. Incomplete or malformed evidence is withheld rather than silently promoted.

§ 02 · Choose by task

Start from the job, not the model.

§ 03 · Evidence

A leaderboard is only as good as the trail behind it.

Measured results, operational claims, and source quality stay separate so a clean-looking table cannot hide a weak comparison.

Every public number carries its audit trail.

CodeSOTA preserves metric direction, the exact benchmark variant, source type, and snapshot used. A vendor claim can be useful, but it is never presented as an independent reproduction.

  • Verified — inspectable evidence meets the publication floor
  • Vendor-reported — primary claim, not independently reproduced
  • Reproduced — an independent run and protocol are available
  • Withheld — the available row does not meet the evidence floor
Read the methodology →
§ 04 · Decision quality

A top score is the start of a decision, not the end.

Benchmark health tells you whether a table still separates models. The decision workspace then adds hardware, openness, language, and evidence constraints.

Matched-snapshot signal
Separating

ARC-AGI-1

31.4-point spread in the current matched frontier snapshot.

Protocol-sensitive

SWE-bench Verified

A strong signal only when harness and scaffold details travel with the score.

Low separation

ARC-Challenge

0.8-point spread across the matched three-model snapshot.

Saturated snapshot

GSM8K

No spread in the matched snapshot; use harder successors for frontier choices.

These are observations from one matched CodeSOTA snapshot, not permanent benchmark properties. Inspect the research question →

§ 05 · Original research

The registry remembers how a result became a claim.

Public experiments, negative results, and lineage connect model comparisons to the research behind them.

Public research graph

Question → experiment → evidence → claim → finding

Follow a decision back through its hypotheses, protocol, evidence, corrections, and negative results. The graph keeps research memory inspectable instead of flattening it into a marketing number.

Explore the research graph →
Questions
7
Experiments
14
Claims
18
Findings
6
LAB → DISTRIBUTION → FACTORY

Find the evidence.
Then try the model.

Research produces the artifact. CodeSOTA helps you inspect the evidence and find a model for your task. For models offered by Fabryka, the API lets you put that choice to work.

Explore models on Fabryka ↗

CodeSOTA is built by Fabryka AI. Hosting availability does not affect benchmark rankings. Model versions and serving configurations may differ from published evaluations.

YOUR NEXT EXPERIMENT
  1. Inspect the result

    Check the model version, task, metric and source.

  2. Confirm availability

    Check the current model catalogue and API limits.

  3. Measure your workload

    Evaluate quality and latency on representative examples.

View the live model catalogue ↗Base URL · https://fabryka.ai/v1
§ 06 · API

The same registry, callable by software.

Agents should not scrape leaderboards. Call a stable, source-aware endpoint and retain the snapshot id so the number in a report can be reproduced later.

  • Public, CORS-open JSON with task-first rankings
  • Benchmark, metric direction, result date, and snapshot id
  • Short task aliases for OCR, code, ASR, TTS, and more
Read the API docs →
GET /api/sotacurl "https://codesota.com/api/sota/ocr?tier=sota"

{
  "task": "ocr",
  "tier": "sota",
  "pick": {
    "model_name": "ovis2-5-9b",
    "score": 63.4,
    "score_metric": "overall-en-private",
    "higher_is_better": true
  },
  "snapshot_id": "2026-04-27"
}
§ 07 · Registry loop

Maintained in public—and open to correction.

Full activity log →
  • ✓ 9,102 results tracked across 163 models

  • + 14 public experiments retain their protocol and outcome

  • ± Superseded claims remain visible in the provenance trail