Design
Which solid notes this experiment tests, and how.
Results
20260901-094622-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag05_jean_zay-f328163
commit: f328163
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 1 |
| condition_0_best_center_eval_accuracy | 0.796875 |
| condition_0_center_eval_accuracy_change | 0.671875 |
| condition_0_center_train_accuracy_change | 0.7142857275903225 |
| condition_0_condition_index | 0 |
| condition_0_final_center_eval_accuracy | 0.796875 |
| condition_0_final_center_eval_parseable_rate | 1 |
| condition_0_final_center_train_accuracy | 0.761904776096344 |
| condition_0_final_center_train_parseable_rate | 1 |
| condition_0_initial_center_eval_accuracy | 0.125 |
| condition_0_initial_center_eval_parseable_rate | 1 |
| condition_0_initial_center_train_accuracy | 0.0476190485060215 |
| condition_0_initial_center_train_parseable_rate | 1 |
| condition_0_lr_direction | 0.25 |
| condition_0_lr_magnitude | 0.5 |
| condition_0_mean_fitness_std | 0.8889708717664083 |
| condition_0_mean_perturbed_train_accuracy | 0.616402147213618 |
| condition_0_mean_perturbed_train_parseable_rate | 1 |
| condition_0_mean_raw_score_std | 0.4688241720199585 |
| condition_0_method_id | 1 |
| condition_0_nonzero_reward_epoch_fraction | 1 |
| condition_0_num_layers | 253 |
| condition_0_rank | 1 |
| condition_0_seed | 0 |
| condition_0_shaped_epoch_fraction | 0 |
| condition_0_sigma_direction | 0.001 |
| condition_0_sigma_magnitude | 0.001 |
| condition_0_tangent_project_direction | 0 |
| decomposed_lr_direction | 0.25 |
| decomposed_lr_magnitude | 0.5 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 4 |
| num_completed_runs | 1 |
| num_epoch_records | 15 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 34146.34162740107 |
Interpretation
Run 20260901-094622-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag05_jean_zay-f328163
is the first completed point in the decomposed Qwen3 fixed-batch calibration
grid. It supports using the low direction/low magnitude scale as a serious
candidate rather than treating the previous decomposed timeout as merely an
implementation-speed problem: the run completed all scheduled epochs, used
exact-answer reward without parseability shaping, and the center-model eval
accuracy improved substantially.
The boring alternative explanation is fixed-batch overfitting. This run cannot
rule that out because it is still one seed and one fixed training batch. The
right next step is to fetch the remaining calibration points, compare whether
this scale is uniquely stable, and then repeat the best decomposed setting on
fresh seeds or batches before making a method-comparison claim.