Design
Which solid notes this experiment tests, and how.
Results
20260901-095021-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag20_jean_zay-fb73f8f
commit: fb73f8f
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 0 |
| decomposed_lr_direction | 0.25 |
| decomposed_lr_magnitude | 2 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 3 |
| num_completed_runs | 0 |
| num_epoch_records | 14 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 33716.349802547134 |
Interpretation
Run 20260901-095021-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag20_jean_zay-fb73f8f
is a partial calibration point. It timed out before producing a completed run
summary, so it should not be treated as a final score for this learning-rate
setting. The flushed epoch and center-eval records are still useful for
diagnosing whether this region of the grid is promising or too slow under the
current evaluation budget.
The boring alternative explanation is runtime rather than learning dynamics:
this configuration may simply need a cheaper final evaluation path to become
comparable with the completed grid points. Use it as partial evidence only
until it is rerun or compared through matched intermediate checkpoints.