Design
Which solid notes this experiment tests, and how.
Results
20260902-082442-qwen3_4b_gsm8k_robustness_plain_bf16_seed1_jean_zay-9765cbb
commit: 9765cbb
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 0 |
| decomposed_lr_direction | 0.25 |
| decomposed_lr_magnitude | 0.5 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 0 |
| method_plain_enabled | 1 |
| noise_reuse | 4 |
| num_center_eval_records | 0 |
| num_completed_runs | 0 |
| num_epoch_records | 0 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| torch_dtype_bfloat16 | 1 |
| torch_dtype_float16 | 0 |
| torch_dtype_float32 | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 36.2086274609901 |
20260902-082442-qwen3_4b_gsm8k_robustness_plain_bf16_seed1_jean_zay-9765cbb
commit: 9765cbb
| metric | value |
|---|---|
| center_eval_do_sample | 0 |
| center_eval_every | 5 |
| center_eval_temperature | 1 |
| center_eval_top_k | 0 |
| completed | 1 |
| condition_0_best_center_eval_accuracy | 0.609375 |
| condition_0_center_eval_accuracy_change | 0.40625 |
| condition_0_center_train_accuracy_change | 0.571428582072258 |
| condition_0_condition_index | 0 |
| condition_0_final_center_eval_accuracy | 0.59375 |
| condition_0_final_center_eval_parseable_rate | 1 |
| condition_0_final_center_train_accuracy | 0.7142857313156128 |
| condition_0_final_center_train_parseable_rate | 1 |
| condition_0_initial_center_eval_accuracy | 0.1875 |
| condition_0_initial_center_eval_parseable_rate | 1 |
| condition_0_initial_center_train_accuracy | 0.1428571492433548 |
| condition_0_initial_center_train_parseable_rate | 1 |
| condition_0_lr | 1 |
| condition_0_mean_fitness_std | 0.8704237858454387 |
| condition_0_mean_perturbed_train_accuracy | 0.6433862785498301 |
| condition_0_mean_perturbed_train_parseable_rate | 1 |
| condition_0_mean_raw_score_std | 0.4666858077049255 |
| condition_0_method_id | 0 |
| condition_0_nonzero_reward_epoch_fraction | 1 |
| condition_0_num_layers | 253 |
| condition_0_rank | 1 |
| condition_0_seed | 1 |
| condition_0_shaped_epoch_fraction | 0 |
| decomposed_lr_direction | 0.25 |
| decomposed_lr_magnitude | 0.5 |
| decomposed_sigma_direction | 0.001 |
| decomposed_sigma_magnitude | 0.001 |
| decomposed_tangent_project_direction | 0 |
| eggroll_lr | 1 |
| eggroll_sigma | 0.001 |
| epochs | 15 |
| eval_batch_size | 8 |
| eval_prompts | 128 |
| generations_per_prompt | 6 |
| max_new_tokens | 1024 |
| method_decomposed_enabled | 0 |
| method_plain_enabled | 1 |
| noise_reuse | 4 |
| num_center_eval_records | 4 |
| num_completed_runs | 1 |
| num_epoch_records | 15 |
| num_methods | 1 |
| num_ranks | 1 |
| num_seeds | 1 |
| num_target_module_patterns | 8 |
| parseability_shaping | 0 |
| population | 126 |
| prompts_per_epoch | 21 |
| temperature | 1 |
| top_k | 0 |
| torch_dtype_bfloat16 | 1 |
| torch_dtype_float16 | 0 |
| torch_dtype_float32 | 0 |
| train_batch_size | 6 |
| wall_clock_seconds | 14487.26922753104 |
Interpretation
Completed after the later fetch. This plain baseline improves over its initial center evaluation but loses ground after the early peak, so it confirms the exact-answer reward signal is usable while also showing instability by the end of the short run.