Design

Which solid notes this experiment tests, and how.

Results

20260901-095518-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag20_jean_zay-c7bc1b0

commit: c7bc1b0

metricvalue
center_eval_do_sample0
center_eval_every5
center_eval_temperature1
center_eval_top_k0
completed0
decomposed_lr_direction0.5
decomposed_lr_magnitude2
decomposed_sigma_direction0.001
decomposed_sigma_magnitude0.001
decomposed_tangent_project_direction0
eggroll_lr1
eggroll_sigma0.001
epochs15
eval_batch_size8
eval_prompts128
generations_per_prompt6
max_new_tokens1024
method_decomposed_enabled1
method_plain_enabled0
noise_reuse4
num_center_eval_records3
num_completed_runs0
num_epoch_records14
num_methods1
num_ranks1
num_seeds1
num_target_module_patterns8
parseability_shaping0
population126
prompts_per_epoch21
temperature1
top_k0
train_batch_size6
wall_clock_seconds33744.030688019935

Interpretation

Run 20260901-095518-qwen3_4b_gsm8k_decomposed_lrdir05_lrmag20_jean_zay-c7bc1b0
is a partial calibration point. It timed out with flushed intermediate records
but no completed-run summary, so its exact table should be read as a diagnostic
trace rather than a final comparison value.

The boring alternative explanation is that this scale is not inherently worse,
only too expensive under the current center-eval cadence. Do not discard it on
timeout alone, but do not pick it over completed settings unless intermediate
records show a clear reason to rerun it.