aqDeveloper environment for foundation models

Steering What Models Learn: Live, Concept-Level Control of Training Runs

See what a model learns as it trains, then steer it at the level of concepts and behaviors. First milestone: 4–6 weeks on fine-tuning.

Why mid-run is blind

Long training runs are mostly launch and hope. The controls we have mid-run are coarse: learning rate, data mix, rollback. None of them let a researcher say "learn this, not that." We shape models almost entirely before training (data, objective) and inspect them after. Between those two points, over thousands of GPU-hours, what the model learns is largely unobserved and uncontrolled.

A debugger for training

The idea is a tool that lets a human see what a model is learning as it trains and steer it with precision, at the level of concepts and behaviors rather than scalar knobs:

  • Watch. Live probes, sparse-autoencoder features, and small capability evals, tracked over steps.
  • Breakpoints. Pause or alert when a feature emerges, or when loss on a slice rises.
  • Hot-reload. Change data weighting, loss terms, or gradient hooks without restarting.
  • Branch. Fork from a checkpoint to test an intervention before committing to it.
  • History. Every intervention is logged as a replayable diff, so the run stays reproducible and auditable.

Controls, from coarse to precise

DataLow, indirect

Reweight examples by concept (e.g. everything that activates feature X)

LossMedium

Hot-add a penalty on a probe's output

GradientHigh

Project a concept direction out of updates; gradient routing to localize knowledge

RepresentationHigh

Preventative activation steering; concept ablation during training (CAFT)

The gradient- and representation-level methods already have early evidence in fine-tuning (gradient routing, Cloud et al. 2024; persona-vector preventative activation steering, 2025; CAFT, 2025). Nobody has yet packaged them as live, composable controls.

First milestone

In 4–6 weeks: a small transformer (~100M parameters) fine-tuned on a task with a known spurious shortcut. A user watches the shortcut feature form, suppresses it live with a gradient or representation intervention, and ends with a model that generalizes correctly out of distribution.

Baselines are no intervention, data-only reweighting, and post-hoc fixes after training. Metrics: out-of-distribution accuracy, in-distribution accuracy retained, compute overhead, and number of interventions needed.

Scope choice: start with fine-tuning, which is short, cheap, and has proven methods. Move to long pretraining once the control primitives and branching harness are solid.

Why it matters

Efficiency: fewer wasted runs, and failures caught early. Control: direct shaping of what models learn, rather than only filtering data. Safety and audit: a replayable record of how a model was shaped, and a way to keep unwanted traits from forming at all instead of removing them later.

The open questions are real. How does a human express "learn the task, not the shortcut"? How early do mid-run signals predict final behavior? How do we measure collateral damage when concepts are entangled? Which interventions help now but hurt the finished model? And how do we hot-swap hooks across a distributed job cheaply and deterministically?

Risks too: observation tools may miss what matters; interventions could hide a behavior rather than prevent it (held-out evals the tool never touches are essential); overhead at scale may limit this to fine-tuning for a while. That is the work.

Not sure if Aquin is right for you?