Design

Which solid notes this experiment tests, and how.

Results

20260901-094918-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag10_jean_zay-b9847c9

commit: b9847c9

metricvalue
center_eval_do_sample0
center_eval_every5
center_eval_temperature1
center_eval_top_k0
completed0
decomposed_lr_direction0.25
decomposed_lr_magnitude1
decomposed_sigma_direction0.001
decomposed_sigma_magnitude0.001
decomposed_tangent_project_direction0
eggroll_lr1
eggroll_sigma0.001
epochs15
eval_batch_size8
eval_prompts128
generations_per_prompt6
max_new_tokens1024
method_decomposed_enabled1
method_plain_enabled0
noise_reuse4
num_center_eval_records3
num_completed_runs0
num_epoch_records15
num_methods1
num_ranks1
num_seeds1
num_target_module_patterns8
parseability_shaping0
population126
prompts_per_epoch21
temperature1
top_k0
train_batch_size6
wall_clock_seconds35821.46464041597

Interpretation

Run 20260901-094918-qwen3_4b_gsm8k_decomposed_lrdir025_lrmag10_jean_zay-b9847c9
is a partial calibration point, not a completed result: Slurm killed it at the
time limit after all epoch records had been flushed but before the final
center evaluation and run summary could complete. The available center evals
still show that this larger magnitude learning rate was not obviously unstable
early; it reached a strong center-model eval score by the last recorded center
checkpoint while using exact-answer reward without parseability shaping.

The boring alternative explanation is that this is another fixed-batch fit
whose apparent eval gain is sensitive to the small deterministic eval slice.
Because the run timed out, it should not be ranked ahead of completed points
without rerunning or reducing evaluation cost. For now it mainly says the
lr_direction=0.25, lr_magnitude=1.0 setting is worth keeping in the
candidate set if the rest of the grid does not dominate it cleanly.