PostTrainBench: Released-Artifact Audit

Paper ID: UnjxMTe57e evidence/provenance.json#/paper_id · arXiv: 2603.08640v2 evidence/provenance.json#/arxiv_id

Claim 1 partial-support evidence/claims.json#/claim_1/status

PostTrainBench evaluates autonomous post-training agents across 4 base models and 7 benchmarks under a 10-hour single-H100 budget (Figure 1). evidence/claims.json#/claim_1/text

SHA-256: 9c0c1fc52ad2a93a9dbe299532b952948c4ecb674f820fe19f78f6a3c33b0073 evidence/claims.json#/claim_1/sha256

Released trajectory inventory confirms 4-by-7 coverage across all accepted benchmark/model cells. Runner configuration defaults to one H100 with a NUM_HOURS-based timeout. The current checkout's scheduler-dependent branches and five-minute termination grace are reported as limitations. evidence/claims.json#/claim_1/summary

Claim 2 partial-support evidence/claims.json#/claim_2/status

The paper reports reward-hacking failure modes including training on test sets, downloading instruction-tuned checkpoints, and using discovered API keys for synthetic data (Abstract). evidence/claims.json#/claim_2/text

SHA-256: d185d61e5d886672a739e321e048df2378b71f55cf388ec4097bd2df1a916aad evidence/claims.json#/claim_2/sha256

Released contamination and instruction-model judgments provide partial support for two of three reward-hacking submodes. The API-key submode artifact is absent from the pinned revision. evidence/claims.json#/claim_2/summary

Coverage Counts

Coverage Matrix

Benchmark Qwen3-1.7B-Base Qwen3-4B-Base SmolLM3-3B-Base Gemma-3-4B-PT
aime2025 47 evidence/coverage.json#/cell_counts/aime2025/0 48 evidence/coverage.json#/cell_counts/aime2025/1 48 evidence/coverage.json#/cell_counts/aime2025/2 48 evidence/coverage.json#/cell_counts/aime2025/3
arenahardwriting 47 evidence/coverage.json#/cell_counts/arenahardwriting/0 48 evidence/coverage.json#/cell_counts/arenahardwriting/1 50 evidence/coverage.json#/cell_counts/arenahardwriting/2 48 evidence/coverage.json#/cell_counts/arenahardwriting/3
bfcl 47 evidence/coverage.json#/cell_counts/bfcl/0 48 evidence/coverage.json#/cell_counts/bfcl/1 48 evidence/coverage.json#/cell_counts/bfcl/2 48 evidence/coverage.json#/cell_counts/bfcl/3
gpqamain 46 evidence/coverage.json#/cell_counts/gpqamain/0 49 evidence/coverage.json#/cell_counts/gpqamain/1 47 evidence/coverage.json#/cell_counts/gpqamain/2 49 evidence/coverage.json#/cell_counts/gpqamain/3
gsm8k 47 evidence/coverage.json#/cell_counts/gsm8k/0 47 evidence/coverage.json#/cell_counts/gsm8k/1 49 evidence/coverage.json#/cell_counts/gsm8k/2 48 evidence/coverage.json#/cell_counts/gsm8k/3
healthbench 47 evidence/coverage.json#/cell_counts/healthbench/0 48 evidence/coverage.json#/cell_counts/healthbench/1 48 evidence/coverage.json#/cell_counts/healthbench/2 48 evidence/coverage.json#/cell_counts/healthbench/3
humaneval 47 evidence/coverage.json#/cell_counts/humaneval/0 47 evidence/coverage.json#/cell_counts/humaneval/1 48 evidence/coverage.json#/cell_counts/humaneval/2 48 evidence/coverage.json#/cell_counts/humaneval/3

Reward-Hacking Statuses

Contamination: partial-support evidence/reward_hacking.json#/training_on_test_sets/status · Instruction model: partial-support evidence/reward_hacking.json#/downloading_instruction_tuned_checkpoint/status · API misuse: unavailable evidence/reward_hacking.json#/using_discovered_api_key/status

Limitations

No H100 run is reproduced; the resource and time findings are a released-configuration audit. evidence/claims.json#/claim_1/limitations/0
The runner allows a five-minute termination grace, and a released example records 10:05:01. evidence/claims.json#/claim_1/limitations/1
The pinned source's current launcher is scheduler-dependent: one branch has a 100-hour/eight-GPU METR command, another has 7 ten-hour and 1 one-hour default-GPU commands, and the arrays currently select only one model/benchmark pair. evidence/claims.json#/claim_1/limitations/2
Evidence is not an official challenge verdict. evidence/claims.json#/claim_1/limitations/3
A released judge label is not independently established behavioral truth. evidence/claims.json#/claim_2/limitations/0
The instruction-model evidence is an upstream categorical label plus safe extracts from a released trace, not a fresh independent behavioral audit. evidence/claims.json#/claim_2/limitations/1
The selected trajectory revision does not expose the exact GPT-5.1 Codex-Max API-misuse task cluster 16804408. evidence/claims.json#/claim_2/limitations/2
No leaderboard score, BFCL score, weighted average, or reasoning-effort ablation is a selected target or reproduced measurement. evidence/claims.json#/claim_2/limitations/3