Finding the Right Job for the Architecture: How I Sidestepped the Generation Penalty in VortexATN
VortexATN held long context but broke text generation. Instead of fixing the bug, I moved the mechanism into an Energy-Based Transformer that only scores Lean 4 tactics and never generates. The ablation is running now.
The short version: I built a filter that helps AI keep its place in long, tangled math proofs. The same filter broke the AI's ability to write. So I stopped asking it to write. I put the filter inside a different kind of model that only reads and grades. That removes the breakage by design. Whether it keeps the benefits is what I am testing right now.
If you've been following my work on AI architecture, you know I've been chasing the "context dilution" problem. Standard transformers, even the ones backed by RoPE, have a real weakness when it comes to deep, structural reasoning: give them a large context window full of complex, multi-branched logic, and their attention degrades into static. They lose the plot.
To work on this, I developed VortexATN (VITX-v2), an attention mechanism that introduces a per-head Phase Bias and a rolling Banded Induction Mask. It acts like a noise-canceling filter, forcing the network to ground itself structurally by analyzing the "directional flux" of the sequence.
In my early synthetic dilution benchmarks, VortexATN earned its keep, holding 80-100% recall through 50 conversational turns while baseline attention collapsed to zero. But it had a cost.
My agent fleet and I documented a flaw I call "Generation Collapse." Because the phase bias rotates the key vectors based on relative frequencies, the model loses its absolute positional grounding during token-by-token generation. It could hold structural integrity across long contexts, but when asked to speak (generate text autoregressively), it would degrade into repeating the End-Of-Sequence (EOS) token.
For a while it looked like a hard trade-off: you could have stable long context, or you could have generation, but not both.
Then I realized I was asking the architecture to do the wrong job.
Energy-Based Verifiers
I am currently deep in formal theorem proving with Lean 4. In formal mathematics, you don't actually want an LLM to guess at text. My Phase 2-fast experiments showed that an LLM cannot reliably evaluate itself: having a base model generate tactics and then asking that same model to re-rank them produced a clean null result (47/374 vs 46/374). The signal has to come from a different inductive bias. You need a system where a candidate generator proposes tactics and a separate, orthogonal verifier scores them.
Enter the Energy-Based Transformer (EBT).
Instead of generating text, an EBT is used discriminatively. It takes a fully-formed (state, tactic) pair, runs a single forward pass at inference time, and returns a scalar "energy" score. Lower energy means the model believes the tactic is compatible and promising according to its structural read, and that score is then handed to the Lean kernel for actual mathematical verification.
Energy-based models never generate text. They read and score. That is the whole loophole.
Vortex-EBT
My coding agents and I opened up the MCMC attention loop inside the AR-EBT backbone, the loop that shapes the energy surface during training, and put VortexATN inside it.
I took an attention mechanism that showed deep structural recall but suffered from generation collapse, and moved it into a verifier architecture that only listens and scores.
The result is the Vortex-EBT.
Mechanically, the generation penalty is isolated by construction. For the Lean 4 state sequence, I route the tensors through VortexATN's phase bias and straight into PyTorch's Scaled Dot-Product Attention (SDPA) backend. For the candidate tactic, I extract the rolling directional flux from the key stream and inject it as a differentiable, banded induction mask directly into the EBT's superdiagonal energy calculation.
This design takes VortexATN's autoregressive failure out of the critical path. Whether it preserves the structural-context advantage is still a hypothesis.
What's Next?
The GPUs in the other room are busy. As of early September, I am running a strict ablation study, training a vanilla EBT baseline against the new Vortex-EBT on the Lean-Mathlib dataset. I will also run synthetic context-dilution sweeps, deliberately contaminating the proof state with up to 32k tokens of irrelevant subgoals, to see whether the vanilla attention collapses while the Vortex-EBT stays flat.
If it holds, it matters beyond Lean 4. The thing I built is a trajectory-energy verifier, and nothing about it is Lean-specific. Picture an agent working through a 100,000-line codebase. A Vortex-EBT could sit in the background, continuously scoring the energy of the agent's trajectory and flagging the moment it drifts from its structural goal.
The ablation is running. When it finishes I will publish the numbers, whether they confirm the hypothesis or hand me another null result.
Dru Edwards