ARM PP PREREGISTRATION — PARAPHRASE PROBE — REGISTERED 2026-08-27, BEFORE SAMPLING

MOTIVATION (responds to a registered limitation, disclosed). Section 9 states that every
arm conditions on a single hash-locked prompt with no paraphrase variant (Sclar et al.,
2024), so the anchor is a property of the (model, prompt string) pair. This arm tests
whether the anchor survives semantic-preserving paraphrase. The paraphrases were authored
by the same pipeline that has seen all prior results; that contamination risk is disclosed
here and bounded by the design: each paraphrase preserves the task's constraints verbatim
in meaning while changing surface form, and all were fixed and hashed before any sampling.

PARAPHRASES (three, fixed; exact strings and SHA-256 in arm_pp_prompts.json, generated by
arm_pp_build.py committed with this file):
  PP-A  reordered clauses, synonym substitutions ("place" for "pack", "touching permitted"),
        constraints as prose rather than a MUST list.
  PP-B  imperative-to-declarative rewrite ("Your task is..."), constraints as numbered list,
        output instruction reworded.
  PP-C  compact single-paragraph statement, mathematical phrasing ("maximize the sum of
        radii subject to containment and pairwise non-overlap").
All three end with the same output-format requirement (raw Python list of [x, y, r]) and
carry the code-free instruction; wrapper A.3 unchanged.

CELLS (two, both discriminating): N = 13 and N = 31. 3 paraphrases x 2 cells x 15 = 90
invocations, same weak-tier subagent channel as arms F/M/CN/CP/RP, runtime rejections
excluded and counted.

SCORING (arm_pp_analysis.py, committed with this file). Arm-F conventions verbatim via
arm_f_repro (parse, validity 1e-6/1e-9, 2e-3 window, mode = most frequent bucket among
valid, ties against the prediction). A (paraphrase, N) cell is EVALUABLE at >= 5 valid.

COMPETING PREDICTIONS.
  P-PP1 (task-level anchoring): the registered prediction (T(4,13) = 1.6250000,
    T(6,31) = 2.5833333) is the modal valid output in >= 4 of the 6 paraphrase-cells.
    Predicted.
  P-PP2 (prompt-string keying): modal in <= 2 of 6, OR pooled validity < 40%. If
    satisfied, S9's (model, prompt string) scoping is PROMOTED from limitation to
    finding, and the abstract's claims are re-scoped to the hash-locked template.
  3 of 6 = PARTIAL, no claim. Fewer than 4 evaluable cells = UNDERPOWERED.
  S-PP1 (strong form): rival-argmax emissions <= 2 across all valid outputs.
  S-PP2 (structure): k = k* majority in >= 4 of 6 evaluable cells.

STOPPING RULE. 90 invocations, one pass, scored once on the complete ledger.

WHAT THIS CANNOT SHOW. Three paraphrases sample the paraphrase space narrowly; a
formatspread-style sweep (Sclar et al., 2024) is out of scope. Nothing here bears on
tiers above the weak tier or on parent-conditioned calls.
