Design

Which solid notes this experiment tests, and how.

Results

20260901-095731-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag05_jean_zay-8a5ad48

commit: 8a5ad48

metricvalue
center_eval_do_sample0
center_eval_every5
center_eval_temperature1
center_eval_top_k0
completed1
condition_0_best_center_eval_accuracy0.7578125
condition_0_center_eval_accuracy_change0.5234375
condition_0_center_train_accuracy_change0.6190476380288601
condition_0_condition_index0
condition_0_final_center_eval_accuracy0.6484375
condition_0_final_center_eval_parseable_rate1
condition_0_final_center_train_accuracy0.6666666865348816
condition_0_final_center_train_parseable_rate1
condition_0_initial_center_eval_accuracy0.125
condition_0_initial_center_eval_parseable_rate1
condition_0_initial_center_train_accuracy0.0476190485060215
condition_0_initial_center_train_parseable_rate1
condition_0_lr_direction1
condition_0_lr_magnitude0.5
condition_0_mean_fitness_std0.8474916815757751
condition_0_mean_perturbed_train_accuracy0.7037037392457326
condition_0_mean_perturbed_train_parseable_rate1
condition_0_mean_raw_score_std0.4432451089223226
condition_0_method_id1
condition_0_nonzero_reward_epoch_fraction1
condition_0_num_layers253
condition_0_rank1
condition_0_seed0
condition_0_shaped_epoch_fraction0
condition_0_sigma_direction0.001
condition_0_sigma_magnitude0.001
condition_0_tangent_project_direction0
decomposed_lr_direction1
decomposed_lr_magnitude0.5
decomposed_sigma_direction0.001
decomposed_sigma_magnitude0.001
decomposed_tangent_project_direction0
eggroll_lr1
eggroll_sigma0.001
epochs15
eval_batch_size8
eval_prompts128
generations_per_prompt6
max_new_tokens1024
method_decomposed_enabled1
method_plain_enabled0
noise_reuse4
num_center_eval_records4
num_completed_runs1
num_epoch_records15
num_methods1
num_ranks1
num_seeds1
num_target_module_patterns8
parseability_shaping0
population126
prompts_per_epoch21
temperature1
top_k0
train_batch_size6
wall_clock_seconds34574.163184434

Interpretation

Run 20260901-095731-qwen3_4b_gsm8k_decomposed_lrdir10_lrmag05_jean_zay-8a5ad48
is a completed calibration point. It finished the full exact-reward schedule,
but the completed-run table should be compared against the lower direction
scales before using it as the next default; a higher direction step may trade
off peak improvement and end-of-run stability.

The boring alternative explanation is fixed-batch sensitivity. This run helps
map the scale surface, but a repeated-seed or fresh-batch run is still required
before interpreting it as genuine reward learning.