Design
Which solid notes this experiment tests, and how.
Results
20260901-095835-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag10_jean_zay-f89db94
commit: f89db94
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 0 |
| decomposed_lr_direction | 1 |
| decomposed_lr_magnitude | 1 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 3 |
| num_completed_runs | 0 |
| num_epoch_records | 15 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 35739.84462727909 |
Interpretation
Run 20260901-095835-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag10_jean_zay-f89db94
is a partial calibration point. It timed out before producing the completed-run
summary, so it should be treated as an intermediate trace only. The setting may
still be informative when compared at matched center checkpoints across the
grid.
The boring alternative explanation is wall-clock budget rather than a failed
learning rate. Do not promote this setting without a completed rerun or a
clear intermediate-checkpoint advantage over completed candidates.