Design

Which solid notes this experiment tests, and how.

Results

20260901-094622-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag05_jean_zay-f328163

commit: f328163

metricvalue
center_eval_do_sample0
center_eval_every5
center_eval_temperature1
center_eval_top_k0
completed1
condition_0_best_center_eval_accuracy0.796875
condition_0_center_eval_accuracy_change0.671875
condition_0_center_train_accuracy_change0.7142857275903225
condition_0_condition_index0
condition_0_final_center_eval_accuracy0.796875
condition_0_final_center_eval_parseable_rate1
condition_0_final_center_train_accuracy0.761904776096344
condition_0_final_center_train_parseable_rate1
condition_0_initial_center_eval_accuracy0.125
condition_0_initial_center_eval_parseable_rate1
condition_0_initial_center_train_accuracy0.0476190485060215
condition_0_initial_center_train_parseable_rate1
condition_0_lr_direction0.25
condition_0_lr_magnitude0.5
condition_0_mean_fitness_std0.8889708717664083
condition_0_mean_perturbed_train_accuracy0.616402147213618
condition_0_mean_perturbed_train_parseable_rate1
condition_0_mean_raw_score_std0.4688241720199585
condition_0_method_id1
condition_0_nonzero_reward_epoch_fraction1
condition_0_num_layers253
condition_0_rank1
condition_0_seed0
condition_0_shaped_epoch_fraction0
condition_0_sigma_direction0.001
condition_0_sigma_magnitude0.001
condition_0_tangent_project_direction0
decomposed_lr_direction0.25
decomposed_lr_magnitude0.5
decomposed_sigma_direction0.001
decomposed_sigma_magnitude0.001
decomposed_tangent_project_direction0
eggroll_lr1
eggroll_sigma0.001
epochs15
eval_batch_size8
eval_prompts128
generations_per_prompt6
max_new_tokens1024
method_decomposed_enabled1
method_plain_enabled0
noise_reuse4
num_center_eval_records4
num_completed_runs1
num_epoch_records15
num_methods1
num_ranks1
num_seeds1
num_target_module_patterns8
parseability_shaping0
population126
prompts_per_epoch21
temperature1
top_k0
train_batch_size6
wall_clock_seconds34146.34162740107

Interpretation

Run 20260901-094622-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag05_jean_zay-f328163
is the first completed point in the decomposed Qwen3 fixed-batch calibration
grid. It supports using the low direction/low magnitude scale as a serious
candidate rather than treating the previous decomposed timeout as merely an
implementation-speed problem: the run completed all scheduled epochs, used
exact-answer reward without parseability shaping, and the center-model eval
accuracy improved substantially.

The boring alternative explanation is fixed-batch overfitting. This run cannot
rule that out because it is still one seed and one fixed training batch. The
right next step is to fetch the remaining calibration points, compare whether
this scale is uniquely stable, and then repeat the best decomposed setting on
fresh seeds or batches before making a method-comparison claim.