Design
Which solid notes this experiment tests, and how.
Results
20260901-095619-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag40_jean_zay-a8f7a80
commit: a8f7a80
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 1 |
| condition_0_best_center_eval_accuracy | 0.765625 |
| condition_0_center_eval_accuracy_change | 0.625 |
| condition_0_center_train_accuracy_change | 0.8571428619325161 |
| condition_0_condition_index | 0 |
| condition_0_final_center_eval_accuracy | 0.75 |
| condition_0_final_center_eval_parseable_rate | 1 |
| condition_0_final_center_train_accuracy | 0.9047619104385376 |
| condition_0_final_center_train_parseable_rate | 1 |
| condition_0_initial_center_eval_accuracy | 0.125 |
| condition_0_initial_center_eval_parseable_rate | 1 |
| condition_0_initial_center_train_accuracy | 0.0476190485060215 |
| condition_0_initial_center_train_parseable_rate | 1 |
| condition_0_lr_direction | 0.5 |
| condition_0_lr_magnitude | 4 |
| condition_0_mean_fitness_std | 0.8956871708234151 |
| condition_0_mean_perturbed_train_accuracy | 0.6978836437066396 |
| condition_0_mean_perturbed_train_parseable_rate | 1 |
| condition_0_mean_raw_score_std | 0.4396172026793162 |
| condition_0_method_id | 1 |
| condition_0_nonzero_reward_epoch_fraction | 1 |
| condition_0_num_layers | 253 |
| condition_0_rank | 1 |
| condition_0_seed | 0 |
| condition_0_shaped_epoch_fraction | 0 |
| condition_0_sigma_direction | 0.001 |
| condition_0_sigma_magnitude | 0.001 |
| condition_0_tangent_project_direction | 0 |
| decomposed_lr_direction | 0.5 |
| decomposed_lr_magnitude | 4 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 4 |
| num_completed_runs | 1 |
| num_epoch_records | 15 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 35548.49023692799 |
Interpretation
Run 20260901-095619-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag40_jean_zay-a8f7a80
is a completed calibration point. It gives a usable end-to-end measurement for
the middle direction scale with the largest magnitude scale in this grid, and
therefore helps bound how aggressive the magnitude update can be before the
next repeated-seed test.
The boring alternative explanation remains fixed-batch fit rather than robust
reward learning. Treat this as one candidate scale, not evidence that the
decomposed parameterisation generalises beyond the fixed batch.