How EvalRank makes evidence-backed decisions

EvalRank turns attributed public benchmark evidence into decisions without hiding comparability limits. It publishes a supported top set when the evidence is sufficient and abstains when it is not.

A benchmark earns a ranking one gate at a time. EvalRank shows exactly how far each capability has moved, and abstains until the last gate is cleared.
  1. 1

    In catalog

    EvalRank understands the question.

  2. 2

    Collecting evidence

    A feed is implemented and parsing.

  3. 3

    Admitted

    Evidence clears rights, identity, and health gates.

  4. 4

    Rankable

    Enough independent families to publish a top set.

Compare like with like

Every published row belongs to one capability cell and one ranking group defined by entity kind, interaction policy, and configuration passport class. Model configurations, agent systems, retrieval components, and arena systems are not flattened onto a shared scale.

An evaluated configuration is identified by its exact passport and revision. Capability evidence is not silently transferred to a different provider offer, harness, scaffold, system prompt, tool set, or quantization.

Preserve uncertainty and honest ties

Intervals and evidence-family coverage stay attached to each ranking. When comparable evidence cannot separate configurations, they remain in the same top set. EvalRank does not force a single winner.

A published top-set claim requires the configured evidence, overlap, and calibration gates. Preview rows disclose the unmet gates and never inherit an active claim.

Bind the decision to the request

A DecisionQuery selects one canonical ranking group and declares the objective and constraints. Cost decisions require an explicit usage profile and can include provider, region, context, budget, cache, and zero-cache sensitivity inputs. Serving-offer economics never leak into capability-only pages.

A DecisionReceipt binds the exact query to one publication snapshot, methodology version, freshness window, reasons, sensitivity checks, and evidence. The outcome is a supported top set or an explicit abstention.

The default share=false request returns a private response without retaining it. Explicit share=true repeats the identical query and retains the content-addressed public-safe receipt so anyone with its URL can replay it.

Treat coverage gaps as repairable

Benchmark health distinguishes active, preview, and unavailable cells. Unavailable means no implemented feed currently supports a public ranking; it is not a declaration that a benchmark or leaderboard is dead.

Feed admission remains source-aware. HTML and linked-data scrapers, including parser recovery work, are valid ways to restore a source when feeds or APIs are incomplete.

Fail closed at the public contract

The seven anonymous public routes use digest-pinned JSON Schemas and OpenAPI definitions. Successful JSON and RFC 9457 Problem Details are validated before the web interface gives them product meaning.

Malformed success documents and non-contract error responses become contract errors. A proxy HTML 404 is never presented as evidence that a receipt or entity does not exist.