Design
Which solid notes this experiment tests, and how.
Results
20260901-095518-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag20_jean_zay-c7bc1b0
commit: c7bc1b0
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 0 |
| decomposed_lr_direction | 0.5 |
| decomposed_lr_magnitude | 2 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 3 |
| num_completed_runs | 0 |
| num_epoch_records | 14 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 33744.030688019935 |
Interpretation
Run 20260901-095518-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag20_jean_zay-c7bc1b0
is a partial calibration point. It timed out with flushed intermediate records
but no completed-run summary, so its exact table should be read as a diagnostic
trace rather than a final comparison value.
The boring alternative explanation is that this scale is not inherently worse,
only too expensive under the current center-eval cadence. Do not discard it on
timeout alone, but do not pick it over completed settings unless intermediate
records show a clear reason to rerun it.