Polinko keeps each result tied to the evidence that produced it.

Each count points back to human-judged pass/fail cases, exclusions, limitations, retained failures, and source notes. The record shows what the result can support from prompt to output.

A count is a pointer into the record.

  • Each count belongs to a named task, case set, and snapshot.
  • A pass means the case held its stated behavioural boundary.
  • A fail can be retained as evidence or evicted from the case set with a recorded reason.
  • Limitations, exclusions, and first-attempt instability remain visible beside the result.
  • Source notes point back to cases, reports, code, or manifests where available.

The current record covers six bounded evidence areas.

OCR baseline

25/25 growth
16/16 focus
0 active fails

Current-set OCR growth and focus cases hold under the tracked replay. Broader image generalisation remains unproven. Open OCR source note.

Co-reasoning

14/14 pass

Style stress cases hold across constraint retention, mode shifts, bounded adaptation, and non-summary collaboration. Open co-reasoning source note.

Retrieval grounding

12/12 retrieval
5/5 file search
0 misses/leaks

Retrieval and file-search cases hold across recall, scoped lookup, and session isolation in the tracked snapshot. Open retrieval source note.

Response behaviour

7/7 pass

Response cases hold across false action claims, uncertainty, invented live verification, directness, roleplay overreach, and pretended memory. One recovered first-attempt failure remains visible. Open response source note.

Hallucination boundary

9/9 pass

Hallucination-boundary cases hold across invented events, motive guesses, archive lore, and fabricated administrative records. Coverage stays limited to selected invention families. Open hallucination source note.

Operator burden

4 pass cases
2 retained fails
1 evicted fail

Operator-burden rows show low-burden control patterns, retained interpretive drift, and one duplicate or noisy case removed from the set. Current evidence is too repetitive for a stronger claim. Open operator-burden source note.

The cases test observable departures from evidence, constraints, and uncertainty.

Source departure

The response invents, misattributes, or claims support absent from the available source material.

Constraint departure

The response ignores an explicit request for directness, exact output, sparse control, or a mid-thread mode change.

Uncertainty departure

The response claims knowledge, memory, verification, or completed action beyond the system's evidence.

Operator burden

The response adds commentary, interpretation, or advice that creates correction work instead of advancing the stated task.

Failures are either retained as evidence or removed from the case set.

Retained failure

The case is valid and the mismatch is in scope. It remains evidence because it reveals a real failure in model behaviour, case design, or method.

Evicted failure

The case is malformed, duplicated, irrelevant, or dominated by setup noise. Removing it protects the evidence set from duplicate pressure.

Eviction is an explicit quality decision with a reviewable reason.

Published evidence is curated from larger working material.

Published evidence includes stable reports, selected case files, manifests, diagrams, and dated research notes. Raw databases, local exports, private transcripts, scratch screenshots, and uncurated evidence dumps enter the record only after explicit review and promotion.

This protects privacy and keeps the public record focused enough to review.

The current evidence supports bounded reliability claims.

  • It supports only the behaviours tested under tracked conditions.
  • The comparison scope is limited to selected models, providers, prompts, and deployment conditions.
  • Transfer to unseen task families remains untested.
  • Human judgement remains visible through the pass/fail gates.
  • Duplicate cases stay grouped instead of counted as independent evidence.

Read the method behind the evidence