ARM CP PREREGISTRATION — REGISTERED 2026-08-27
Perturbed-container probe: same difficulty, destroyed lexical overlap

REGISTRATION BLOCK (added at registration; the draft body below is unchanged).
- Drafted 2026-08-22 as arm_cp_preregistration_DRAFT.txt (committed); registered
  2026-08-27 by removing the DRAFT marker, generating prompt hashes, and committing
  arm_cp_build.py + arm_cp_analysis.py + arm_cp_prompts.json alongside this file,
  pushed BEFORE the first invocation.
- Prompt SHA-256 per cell: in arm_cp_prompts.json (arm_cp_build.py regenerates and
  self-checks against the draft's stated N=13 and N=58 mapped values).
- Sampling channel: the same weak-tier subagent channel as arms F/M/CN (Claude Code
  Task subagent, claude-3-5-haiku class weak tier, dispatch wrapper A.3 verbatim,
  zero-shot, no tools, no code). n = 15 per cell, 75 invocations, one pass.
- All five registered cells are discriminating (prediction != family argmax at > 2e-3),
  confirmed by arm_cp_build.py output at registration time.
- No sampling occurred before the registration commit.

MOTIVATION (post-hoc, disclosed). Arm CN showed the closed form predicts the mode at N absent
from every cited scoreboard, which bears on recall of PUBLISHED VALUES. It does not bear on
recall of the TASK TEXT: every arm so far uses the byte-identical bare template A.1 ("unit
square [0,1]x[0,1]"), the phrasing closest to the benchmark's own. A reviewer can hold that the
template is keyed to that string. This arm keeps the geometry and the optimum structure
invariant while changing every lexical handle.

PERTURBATION (one, fixed). Container becomes the square [3, 5] x [3, 5] (side 2, offset 3).
Under the similarity map x -> 2x + 3 every valid packing of the unit square maps to a valid
packing of the new container with each radius doubled and the sum of radii doubled; the
optimum, the recipe family, the branch rule and the trap zones are invariant. Predicted
values are therefore exactly 2 * V(k*, m) or 2 * T(k*, N); the model never sees "[0,1]",
"unit square", or the published magnitudes (2.63... becomes 5.27...). Radii and coordinates
are scored after mapping back to the unit square (x -> (x - 3) / 2, r -> r / 2), so the
registered tolerances (1e-6 primary, 1e-9 logged) and the 2e-3 value window apply unchanged
in unit-square units.

CELLS (five): the three discriminating original-arm cells and two held-out discriminating
cells, so every cell can separate anchoring from family search.
  N=13  k*=4  T(4,13) -> 2 * 1.6250000 = 3.2500000   rival 2 * 1.7761424
  N=21  k*=5  T(5,21) -> 2 * 2.1000000 = 4.2000000   rival 2 * 2.2588835
  N=31  k*=6  T(6,31) -> 2 * 2.5833333 = 5.1666667   rival 2 * 2.7485281
  N=58  k*=8  T(8,58) -> 2 * 3.6250000 = 7.2500000   rival 2 * 3.7662801
  N=75  k*=9  T(9,75) -> 2 * 4.1666667 = 8.3333333   rival 2 * 4.2847718
(Values to be regenerated by arm_cp_build.py from the arm-M V/T functions, self-checked against
the registered N=13 and N=58 values above; the rival column is the family argmax at k*-1.)

PROPOSER AND PROTOCOL. Identical to arm CN: weak-tier subagent channel, bare prompt, zero-shot,
code-free, no tools, dispatch wrapper A.3 verbatim, n = 15 per cell, 75 invocations, runtime
rejections excluded and counted, every completed message logged verbatim and scored.

PROMPT. Template A.1 with two substitutions only: "unit square [0,1]x[0,1]" -> "square
[3,5]x[3,5]" (both occurrences) and the containment line rewritten with the new bounds
(x-r >= 3, x+r <= 5, same for y). No other token changes. SHA-256 per cell to be written to
arm_cp_prompts.json before sampling.

SCORING (arm_cp_analysis.py, to be committed before sampling). Map each emitted circle back
to the unit square, then apply arm-F conventions unchanged (parse, validity, 2e-3 window,
mode = most frequent bucket among valid, ties count against the prediction, structural k from
dominant radius after mapping). EVALUABLE at >= 5 valid per cell; otherwise UNDERPOWERED.

COMPETING PREDICTIONS.
  P-CP1 (construction, lexically independent): modal valid output equals the mapped
    prediction at >= 4 of 5 evaluable cells. Predicted.
  P-CP2 (lexical keying): anchoring is tied to the "[0,1]" phrasing; modal output equals the
    mapped prediction at <= 2 of 5 evaluable cells, OR validity collapses below 40% pooled
    (the model cannot execute the grid in an offset frame).
  3 of 5 = PARTIAL; < 4 evaluable cells = UNDERPOWERED; no branch claimed in either case.
  S-CP1 (strong form): rival-argmax emissions <= 1 across all discriminating-cell validities.
  S-CP2 (structure): k = k* majority at >= 4 of 5 cells.
  FALSIFIER F-CP1: P-CP2 satisfied. Consequence: Section 8's contamination paragraph is
    rewritten to state that the anchor is keyed to the benchmark phrasing, and the abstract's
    "closed form" claim is rescoped to the literal task text.

STOPPING RULE. 75 invocations, one pass, no top-ups. Analysis script run once on the complete
ledger; the frozen report is the only result reported.

WHAT THIS ARM CANNOT SHOW. Absence of the unit-square task from pretraining; anything about
parent-conditioned calls; anything about tiers above the weak tier.
