Evidence ledger

What has actually been observed?

Palimpsestus separates structural conformance, controlled behavioral evidence, official-evaluation milestones, null results, and unresolved attribution. A result is not promoted beyond the evidence that warrants it.

Architecture conformance
192/192
semantic obligations in the audited ARC-AGI-2 realization
160/160
derivation-ancestry edges
ObservedStructural
ARC-AGI-3 · ls20
0 / 250
repeated disproven actions with provenance-faithful recurrence
35 / 250
matched learner, 14%
ObservedPreregistered
Official ARC-AGI-3 pathway
0.08
official Kaggle publicScore from a completed submission
ObservedAttribution unresolved

ARC-AGI-2

The frozen realization conformed to the extracted semantic and ancestry obligations without modifying C₁. Predictive Fold geometry received confirmatory support.

Null result: no measurable controller endpoint advantage was observed.

Still unresolved: preservation of useful operational horizon under finite budget.

ARC-AGI-3 controlled Crucible

On preregistered official ls20, provenance-faithful recurrence reduced repeated disproven actions from 35/250 in the matched learner to 0/250.

Control: su15 showed no measurable behavioral separation.

Capability boundary: neither architecture achieved level progress in that controlled experiment. Goal-directed utility therefore remains unresolved.

Official evaluation

The non-zero score is real. Its ancestry is not yet closed.

An official Kaggle ARC-AGI-3 submission reached COMPLETE status and returned publicScore = 0.08. The sealed record does not yet establish that the score descended specifically from Palimpsestus agent execution on hidden tasks.

Claim boundary

Warranted: a completed official submission returned a non-zero public score of 0.08.

Not yet warranted: “Palimpsestus scored 0.08 on hidden ARC-AGI-3 tasks” or “Palimpsestus solved 8% of ARC-AGI-3.”

Next evidence

The next result must discriminate.

The next major experiment is intended to compare Palimpsestus against Active Inference, NARS/AERA, and a matched conventional learner in environments where the architectures should make different predictions before execution.

Read the research agenda →