Design
Which solid notes this experiment tests, and how.
Results
20260901-100051-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag40_jean_zay-fa7d6c8
commit: fa7d6c8
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 0 |
| decomposed_lr_direction | 1 |
| decomposed_lr_magnitude | 4 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 3 |
| num_completed_runs | 0 |
| num_epoch_records | 13 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 34256.93529732688 |
Interpretation
Run 20260901-100051-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag40_jean_zay-fa7d6c8
is a partial calibration point. It timed out before the final summary and is
therefore not directly comparable with completed grid points. Its intermediate
records are still useful for checking whether the most aggressive direction
and magnitude combination was plainly unstable.
The boring alternative explanation is that the missing final result is a
runtime/evaluation-cost issue rather than a learning failure. Keep this as a
diagnostic boundary case, not as a candidate scale unless a cheaper rerun is
explicitly needed.