ARM RP PREREGISTRATION — DRAFT, NOT REGISTERED, NO SAMPLING AUTHORISED
Direct-recall probe: can the proposer STATE the published values it is accused of recalling?
Drafted 2026-08-22 for the TMLR review stage of the anchoring paper. Becomes a registration
only when the DRAFT marker is removed, prompt hashes are written to arm_rp_prompts.json,
arm_rp_analysis.py is committed alongside, and the commit is pushed before the first
invocation. Until then this file claims nothing.

MOTIVATION (post-hoc, disclosed). Section 8 names a "canary-string test" as unrun. The classic
canary design (a unique string planted in a benchmark, later elicited from the model) cannot be
applied retrospectively to a benchmark we did not author. The nearest applicable test asks the
recall question directly: if the weak-tier proposer's template values come from memory of
published packings, it should be able to state those published values when asked for them
outright — the easiest possible recall task. If instead it returns family values V/T, returns
nothing, or returns values unrelated to the scoreboard, the recall reading loses its simplest
prediction. This is a necessary-condition probe, not proof of absence from pretraining.

CELLS. Six N, three with a published sum-of-radii value in the cited scoreboard and three
without:
  scoreboard:  N=26 (2.6358...; AlphaEvolve/ShinkaEvolve/HELIX cluster), N=32 (scoreboard
               value as cited in Section 7), N=13 (scoreboard value as cited in Section 7)
  held-out:    N=50, N=62, N=75 (no published sum-of-radii value; arm CN cells)
Exact scoreboard values and their sources are copied into arm_rp_build.py from Section 7's
citations before registration, with the citation for each.

PROPOSER AND PROTOCOL. Same weak-tier channel as arms F/M/CN, no tools, dispatch wrapper A.3,
n = 15 per cell, 90 invocations, runtime rejections excluded and counted, every completed
message logged verbatim.

PROMPT (new, to be hashed). "What is the best-known maximum sum of radii for {n}
non-overlapping circles packed inside the unit square [0,1]x[0,1]? Answer with the number
only, to at least six decimal places, and nothing else. If you do not know, answer UNKNOWN."

SCORING (arm_rp_analysis.py). A response RECALLS if it parses as a number within 2e-3 of the
published value for that N. A response is FAMILY if it is within 2e-3 of V(k*, m) or T(k*, N)
for that N. UNKNOWN and unparseable responses are counted separately. No geometry is emitted,
so validity does not apply.

COMPETING PREDICTIONS.
  P-RP1 (construction): RECALLS in <= 2 of 45 scoreboard-cell responses, and FAMILY responses
    at scoreboard cells do not exceed FAMILY responses at held-out cells by more than 20
    percentage points (the family value is a computation, not a memory, so it should not be
    easier at published N). Predicted.
  P-RP2 (recall): RECALLS in >= 10 of 45 scoreboard-cell responses (>= 22%), with RECALLS at
    held-out cells necessarily zero (no value exists); this is the asymmetry memory produces.
  Between 3 and 9 RECALLS = PARTIAL, reported with no claim.
  FALSIFIER F-RP1: P-RP2 satisfied. Consequence: Section 8 states that the weak tier can
    recall scoreboard values on request, and the construction reading of arm CN is weakened to
    "construction at held-out N, recall not excluded at published N".

STOPPING RULE. 90 invocations, one pass. Analysis run once on the complete ledger.

WHAT THIS ARM CANNOT SHOW. A model that fails to state a value may still have been shaped by
it; absence of recall on request is not absence from pretraining. The arm bounds the simplest
recall mechanism only.
