Design
Which solid notes this experiment tests, and how.
Results
20260901-095408-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag10_jean_zay-b2aca31
commit: b2aca31
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 1 |
| condition_0_best_center_eval_accuracy | 0.765625 |
| condition_0_center_eval_accuracy_change | 0.640625 |
| condition_0_center_train_accuracy_change | 0.7619047723710537 |
| condition_0_condition_index | 0 |
| condition_0_final_center_eval_accuracy | 0.765625 |
| condition_0_final_center_eval_parseable_rate | 1 |
| condition_0_final_center_train_accuracy | 0.8095238208770752 |
| condition_0_final_center_train_parseable_rate | 1 |
| condition_0_initial_center_eval_accuracy | 0.125 |
| condition_0_initial_center_eval_parseable_rate | 1 |
| condition_0_initial_center_train_accuracy | 0.0476190485060215 |
| condition_0_initial_center_train_parseable_rate | 1 |
| condition_0_lr_direction | 0.5 |
| condition_0_lr_magnitude | 1 |
| condition_0_mean_fitness_std | 0.9046175559361775 |
| condition_0_mean_perturbed_train_accuracy | 0.6708995143572489 |
| condition_0_mean_perturbed_train_parseable_rate | 1 |
| condition_0_mean_raw_score_std | 0.4514849344889323 |
| condition_0_method_id | 1 |
| condition_0_nonzero_reward_epoch_fraction | 1 |
| condition_0_num_layers | 253 |
| condition_0_rank | 1 |
| condition_0_seed | 0 |
| condition_0_shaped_epoch_fraction | 0 |
| condition_0_sigma_direction | 0.001 |
| condition_0_sigma_magnitude | 0.001 |
| condition_0_tangent_project_direction | 0 |
| decomposed_lr_direction | 0.5 |
| decomposed_lr_magnitude | 1 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 1 |
| method_plain_enabled | 0 |
| noise_reuse | 4 |
| num_center_eval_records | 4 |
| num_completed_runs | 1 |
| num_epoch_records | 15 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 35730.995360302 |
Interpretation
Run 20260901-095408-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag10_jean_zay-b2aca31
is a completed calibration point. It finished the planned epoch and center-eval
schedule under exact-answer reward only, so it belongs in the shortlist of
usable decomposed scales for a repeated-seed or fresh-batch follow-up.
The boring alternative explanation is fixed-batch overfitting. This run does
not rule that out, because the training prompts and deterministic eval slice
are still small. It should be compared with the other completed points first,
then repeated outside this fixed batch before supporting a method claim.