Paper ID: UnjxMTe57e evidence/provenance.json#/paper_id · arXiv: 2603.08640v2 evidence/provenance.json#/arxiv_id
PostTrainBench evaluates autonomous post-training agents across 4 base models and 7 benchmarks under a 10-hour single-H100 budget (Figure 1). evidence/claims.json#/claim_1/text
SHA-256: 9c0c1fc52ad2a93a9dbe299532b952948c4ecb674f820fe19f78f6a3c33b0073 evidence/claims.json#/claim_1/sha256
Released trajectory inventory confirms 4-by-7 coverage across all accepted benchmark/model cells. Runner configuration defaults to one H100 with a NUM_HOURS-based timeout. The current checkout's scheduler-dependent branches and five-minute termination grace are reported as limitations. evidence/claims.json#/claim_1/summary
The paper reports reward-hacking failure modes including training on test sets, downloading instruction-tuned checkpoints, and using discovered API keys for synthetic data (Abstract). evidence/claims.json#/claim_2/text
SHA-256: d185d61e5d886672a739e321e048df2378b71f55cf388ec4097bd2df1a916aad evidence/claims.json#/claim_2/sha256
Released contamination and instruction-model judgments provide partial support for two of three reward-hacking submodes. The API-key submode artifact is absent from the pinned revision. evidence/claims.json#/claim_2/summary
viewer_data evidence/coverage.json#/excluded_auxiliary_data/top_level_path; files: 2397 evidence/coverage.json#/excluded_auxiliary_data/file_count; directories: 1 evidence/coverage.json#/excluded_auxiliary_data/directory_count; counted as task root: false evidence/coverage.json#/excluded_auxiliary_data/counted_as_task_root| Benchmark | Qwen3-1.7B-Base | Qwen3-4B-Base | SmolLM3-3B-Base | Gemma-3-4B-PT |
|---|---|---|---|---|
aime2025 |
47 evidence/coverage.json#/cell_counts/aime2025/0 | 48 evidence/coverage.json#/cell_counts/aime2025/1 | 48 evidence/coverage.json#/cell_counts/aime2025/2 | 48 evidence/coverage.json#/cell_counts/aime2025/3 |
arenahardwriting |
47 evidence/coverage.json#/cell_counts/arenahardwriting/0 | 48 evidence/coverage.json#/cell_counts/arenahardwriting/1 | 50 evidence/coverage.json#/cell_counts/arenahardwriting/2 | 48 evidence/coverage.json#/cell_counts/arenahardwriting/3 |
bfcl |
47 evidence/coverage.json#/cell_counts/bfcl/0 | 48 evidence/coverage.json#/cell_counts/bfcl/1 | 48 evidence/coverage.json#/cell_counts/bfcl/2 | 48 evidence/coverage.json#/cell_counts/bfcl/3 |
gpqamain |
46 evidence/coverage.json#/cell_counts/gpqamain/0 | 49 evidence/coverage.json#/cell_counts/gpqamain/1 | 47 evidence/coverage.json#/cell_counts/gpqamain/2 | 49 evidence/coverage.json#/cell_counts/gpqamain/3 |
gsm8k |
47 evidence/coverage.json#/cell_counts/gsm8k/0 | 47 evidence/coverage.json#/cell_counts/gsm8k/1 | 49 evidence/coverage.json#/cell_counts/gsm8k/2 | 48 evidence/coverage.json#/cell_counts/gsm8k/3 |
healthbench |
47 evidence/coverage.json#/cell_counts/healthbench/0 | 48 evidence/coverage.json#/cell_counts/healthbench/1 | 48 evidence/coverage.json#/cell_counts/healthbench/2 | 48 evidence/coverage.json#/cell_counts/healthbench/3 |
humaneval |
47 evidence/coverage.json#/cell_counts/humaneval/0 | 47 evidence/coverage.json#/cell_counts/humaneval/1 | 48 evidence/coverage.json#/cell_counts/humaneval/2 | 48 evidence/coverage.json#/cell_counts/humaneval/3 |
Contamination: partial-support evidence/reward_hacking.json#/training_on_test_sets/status · Instruction model: partial-support evidence/reward_hacking.json#/downloading_instruction_tuned_checkpoint/status · API misuse: unavailable evidence/reward_hacking.json#/using_discovered_api_key/status