Design
Which solid notes this experiment tests, and how.
Results
20260901-095731-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag05_jean_zay-8a5ad48
commit: 8a5ad48
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 1 |
| condition_0_best_center_eval_accuracy | 0.7578125 |
| condition_0_center_eval_accuracy_change | 0.5234375 |
| condition_0_center_train_accuracy_change | 0.6190476380288601 |
| condition_0_condition_index | 0 |
| condition_0_final_center_eval_accuracy | 0.6484375 |
| condition_0_final_center_eval_parseable_rate | 1 |
| condition_0_final_center_train_accuracy | 0.6666666865348816 |
| condition_0_final_center_train_parseable_rate | 1 |
| condition_0_initial_center_eval_accuracy | 0.125 |
| condition_0_initial_center_eval_parseable_rate | 1 |
| condition_0_initial_center_train_accuracy | 0.0476190485060215 |
| condition_0_initial_center_train_parseable_rate | 1 |
| condition_0_lr_direction | 1 |
| condition_0_lr_magnitude | 0.5 |
| condition_0_mean_fitness_std | 0.8474916815757751 |
| condition_0_mean_perturbed_train_accuracy | 0.7037037392457326 |
| condition_0_mean_perturbed_train_parseable_rate | 1 |
| condition_0_mean_raw_score_std | 0.4432451089223226 |
| condition_0_method_id | 1 |
| condition_0_nonzero_reward_epoch_fraction | 1 |
| condition_0_num_layers | 253 |
| condition_0_rank | 1 |
| condition_0_seed | 0 |
| condition_0_shaped_epoch_fraction | 0 |
| condition_0_sigma_direction | 0.001 |
| condition_0_sigma_magnitude | 0.001 |
| condition_0_tangent_project_direction | 0 |
| decomposed_lr_direction | 1 |
| decomposed_lr_magnitude | 0.5 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 4 |
| num_completed_runs | 1 |
| num_epoch_records | 15 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 34574.163184434 |
Interpretation
Run 20260901-095731-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag05_jean_zay-8a5ad48
is a completed calibration point. It finished the full exact-reward schedule,
but the completed-run table should be compared against the lower direction
scales before using it as the next default; a higher direction step may trade
off peak improvement and end-of-run stability.
The boring alternative explanation is fixed-batch sensitivity. This run helps
map the scale surface, but a repeated-seed or fresh-batch run is still required
before interpreting it as genuine reward learning.