Polinko evaluates AI behaviour through evidence, constraints, and retained failures.
Polinko uses binary evals, signal traces, and retained failures to inspect how model output is generated and transformed without mistaking coherence for reliability.
Coherence alone does not make an answer reliable.
An answer may sound complete while inventing a source, ignoring an explicit constraint, overstating what was verified, or adding interpretation that makes direct work harder. Polinko keeps those departures visible as inspectable cases.
The project asks what becomes visible when failure, rather than pass rate, is treated as the main research signal.
Six investigations are at different stages of maturity.
Each investigation has its own task shape, current read, and evidence boundary. Results stay within the lane that produced them.
OCR generalisation
The current image set is stable. The next test covers new images and harder visual conditions.
Stable baselineCo-reasoning
Tests constraint retention, mode shifts, bounded adaptation, and collaboration beyond summary.
Promoted laneRetrieval grounding
Tests recall and scoped search with explicit miss, isolation, and leak checks.
OperationalisedResponse behaviour
Tests false action claims, unsupported certainty, invented live verification, and pretended memory.
OperationalisedHallucination boundaries
Tests specific unsupported inventions, including archive lore and fabricated records.
OperationalisedOperator burden
Tracks when unwanted commentary, interpretation, or advice adds work instead of helping.
Thin laneEvidence moves from source artefact to bounded claim.
I define the questions and tests, direct the collaborative work, contribute to the engineering, and decide whether the evidence can support a public claim.
The current baseline spans six evaluation surfaces.
Each count is shown with the task, cases, limitations, and source notes that define what it can support.
- OCR baseline
- 25/25
- Co-reasoning
- 14/14
- Retrieval
- 12/12 + 5/5
- Response behaviour
- 7/7
- Hallucination boundary
- 9/9
- Operator burden
- 4 / 2 / 1
Growth stability, with 16/16 focus stability and no active fail-history cases.
Tracked style stress cases across seven collaboration behaviours.
Retrieval and file-search cases with zero tracked miss or leak in the snapshot.
The tracked run closed cleanly, with one recovered first-attempt failure still visible.
Nine tracked invention cases across distinct claim families.
Four pass anchors, two retained failures, and one evicted duplicate or noisy case.
Results apply to the cases and conditions tested.
A passing evaluation supports a finding within its named cases, method, and evidence conditions.
Limitations, exclusions, retained failures, and unresolved questions remain visible alongside the result.
I lead the research and make the final judgements.
I define the research questions, evaluation boundaries, evidence criteria, interpretations, and final claims. I contribute to the engineering and work with AI systems on dialogue, implementation, analysis, stress-testing, and documentation. The process is collaborative, while source authority and publication decisions are mine.
Beta 2.3 is the current evidence baseline.
- Beta 2.3 preserves the evidence boundary for the current method work.
- OCR is moving from current-set stability into broader generalisation pressure.
- Pre-Beta 2.4 keeps row and case evidence visible before lane-level claims.
- New evidence folders appear only when evidence is ready to promote.
- Historical hypotheses that failed remain visible in the record.