Design
Which solid notes this experiment tests, and how.
Results
20260901-095158-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag40_jean_zay-ab9bbb8
commit: ab9bbb8
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 0 |
| decomposed_lr_direction | 0.25 |
| decomposed_lr_magnitude | 4 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 3 |
| num_completed_runs | 0 |
| num_epoch_records | 14 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 34026.89495655312 |
Interpretation
Run 20260901-095158-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag40_jean_zay-ab9bbb8
is a partial calibration point. It timed out before the final center-model
evaluation and completed-run summary, so the setting should not be ranked
against completed runs on final accuracy. The partial records are enough to
check whether the high magnitude step showed early instability under exact
reward, but not enough for a selection decision by themselves.
The boring alternative explanation is that the fixed batch and expensive
evaluation schedule, rather than this scale choice, caused the missing final
summary. Keep this point in the diagnostic set, but prefer completed runs when
choosing the next repeated-seed test.