Design
Which solid notes this experiment tests, and how.
Results
20260901-095300-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag05_jean_zay-5a44d6a
commit: 5a44d6a
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 0 |
| decomposed_lr_direction | 0.5 |
| decomposed_lr_magnitude | 0.5 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 3 |
| num_completed_runs | 0 |
| num_epoch_records | 15 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 35373.568909072084 |
Interpretation
Run 20260901-095300-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag05_jean_zay-5a44d6a
is a partial calibration point. It reached the epoch loop and flushed center
checks, but it did not complete the final summary before the time limit. That
makes it useful for trend inspection, not for declaring this scale better or
worse than completed settings.
The boring alternative explanation is again runtime pressure from the
evaluation protocol. This setting should be compared at matched intermediate
checkpoints or rerun with cheaper evaluation before it is used to choose the
decomposed method scale.