AlphaFold and Boltz: structure prediction and recycling

I’ve been trying to understand how AlphaFold works, and writing this down seemed like a good way to get my thoughts a little clearer.

The focus is the history, what changed architecturally from one generation to the next, and how those changes carry through to Boltz-2. I also want to spend some time on recycling, which I think was ahead of its time as a practical use of repeated inference: AlphaFold 2 used its own intermediate representations and predicted geometry to improve the next pass. That makes it an interesting example to compare with today’s looped transformers.

Sequence and structure

Protein structure prediction requires reasoning about spatial relationships that sequence order does not directly encode. Residues far apart in sequence can be adjacent in the folded structure, so the model needs representations of both individual residues and their relationships.

A sequence feeds a pair representation and a folded chain with contacts between sequence-distant residues.
Figure 1. Sequence-distant residues can be neighbors in space.

AlphaFold predicts a structure; it does not simulate the physical process of folding. Its intermediate structures are computational hypotheses, not snapshots of the molecule moving through time. The distinction matters when interpreting recycling or diffusion steps.

From AlphaFold 1 to Boltz-2

Year Model Main architectural change
2018 · CASP13 AlphaFold 1 Convolutional distance prediction defines a learned potential; separate optimization constructs the coordinates.
2020 · CASP14 / 2021 paper AlphaFold 2 Evoformer and a learned structure module integrate relationship reasoning and coordinate construction; recycling feeds predictions into another pass.
2024 AlphaFold 3 Pairformer representations condition atom-coordinate diffusion, expanding prediction to more kinds of molecular complexes.
2025 Boltz-2 An open model in the AlphaFold 3 family adds binding-affinity prediction alongside structure generation.

This history involves several overlapping ideas. A graph specifies which entities and relationships are represented. Attention controls how information is exchanged. Diffusion specifies a way to generate an output. AlphaFold 3 combines pairwise graph reasoning, attention, and diffusion in the same system. Its use of diffusion does not mean that attention or transformers failed.

Evolutionary evidence and pair representations

AlphaFold 2 uses a multiple sequence alignment, or MSA, of related protein sequences. Each row is a related sequence, and each column corresponds to an aligned position. Experimentally determined structures of related proteins can also enter as templates when available.

Evolution supplies clues about interactions. If two positions repeatedly change together, a change at one may require a compensating change at the other to maintain a useful structure or function. These correlations are informative, but shared ancestry and indirect interactions also create them. Correlated changes are not proof of direct contact.

Alongside the MSA, the network maintains a pair representation. For N residues, this is an N × N grid of feature vectors. The vector zᵢⱼ stores learned information about the relationship between residues i and j. It can encode much more than a distance or a contact probability.

Aligned sequences provide evidence for a grid of residue-pair features.
Figure 2. Evolutionary patterns inform residue-pair features.

The Evoformer updates these two representations together. Attention moves information across aligned positions and related sequences. An outer-product update transfers information from the MSA into the pair grid. Pair features also bias attention in the MSA, so current relationship estimates influence how sequence evidence is processed.

The pair grid keeps a dedicated feature vector for each residue pair. Structural training data teaches the model which patterns are useful for predicting geometry; the architecture organizes the computation around those patterns.

Triangle updates and graph reasoning

Treat residues as nodes in a dense graph and pair vectors as edge features. An important part of AlphaFold’s computation updates an edge using other edges.

To revise the relationship between i and j, consider a third residue k. The relationships i–k and j–k constrain what is plausible for i–j. If two residues are each close to the same third residue, their relationship cannot be chosen independently of that information.

Edges through a third residue k inform the update to the relationship between i and j.
Figure 3. A pair is updated using the other two edges of a triangle.

Triangle multiplication combines transformed pair features through the third residue. Triangle attention weights relevant relationships around a triangle. These are learned operations on feature vectors. They encourage compatible relationships without explicitly enforcing geometric rules such as the triangle inequality.

A collection of plausible pairwise predictions may still be impossible to realize as one three-dimensional object. Triangle updates allow information about one relationship to revise another. The distinctive graph idea is therefore reasoning over relationships between relationships, not simply passing messages between residue nodes.

From pair features to coordinates

AlphaFold 2’s structure module assigns each residue a local coordinate frame: a position and an orientation. It updates these frames using invariant point attention (IPA).

Invariant point attention combines feature similarity, pair information, and distances between learned points attached to the local frames. The points are transformed into a common coordinate system so their relative geometry can affect attention.

Residue frames and learned points retain their relative geometry under a global rotation and translation.
Figure 4. Attention is invariant to a global rigid motion; coordinate updates are equivariant.

Rotating or translating the whole molecule leaves those point distances unchanged. The scalar attention weights are therefore invariant to a global rigid motion. The coordinate updates are equivariant: they rotate and translate consistently with the molecule. This lets the network use geometry without making its answer depend on an arbitrary orientation in space.

Together, the MSA, pair representation, triangle updates, and structure module connect evolutionary evidence to geometry while respecting rotation and translation symmetry.

Recycling in AlphaFold 2

Recycling runs the same model again using information from its previous pass. AlphaFold 2 carries forward pair features, features from the MSA’s target-sequence row, and geometry derived from predicted coordinates. The original input evidence remains available.

Parameters are shared across cycles, although the Evoformer blocks within each cycle have their own parameters. Recycling adds inference computation; it does not update the model’s weights.

Input evidence enters a shared Evoformer and structure module while previous representations and coordinates return through a feedback loop.
Figure 5. Previous features and predicted geometry inform the next cycle.

Constructing coordinates reveals the arrangement implied by the current pair features. Feeding that geometry back gives the next pass context for revising those relationships. A tentative placement of two regions, for example, can inform another round of pair updates. This is a useful interpretation of the feedback, rather than an explicit constraint-solving algorithm.

The next cycle also builds on earlier work instead of reconstructing every relationship from the input alone. The authors’ ablations found recycling important for accuracy.

AlphaFold 3 and diffusion

AlphaFold 3 predicts complexes involving proteins, DNA, RNA, small molecules, ions, and modified residues. Its main representation trunk, the Pairformer, carries single-token and pair features. An earlier module processes the MSA; the main trunk no longer maintains AlphaFold 2’s persistent MSA representation. Triangle operations and attention remain central.

The larger change is coordinate generation. A diffusion module produces atom coordinates conditioned on the trunk’s representations. During training, the denoiser learns to recover clean coordinates from corrupted ones. During inference, a sampler starts with noisy coordinates and repeatedly applies denoising updates under a changing noise schedule.

Recycled Pairformer representations condition a separate sequence of atom-coordinate denoising steps.
Figure 6. Representation recycling precedes coordinate sampling.

There are two different repeated computations here. AlphaFold 3 first recycles single and pair representations through the trunk. It then samples coordinates using the completed trunk as conditioning. Sampled coordinates are not fed back through the trunk as they are in AlphaFold 2.

Thoughts on diffusion

I think diffusion works well here because it gives the model repeated chances to adjust the whole structure. Molecular constraints are coupled: fitting a ligand can require changes in the surrounding pocket. A basic atom-by-atom autoregressive decoder fixes earlier coordinates unless it adds revision or search. Denoising lets those placements keep changing together.

The noise levels also seem useful as a way to organize what the model learns. Heavy noise emphasizes the overall arrangement; slight noise emphasizes local stereochemistry. That gives the model practice correcting geometry at both scales, instead of requiring one update to get everything right.

The sampling side makes sense to me too. Uncertainty about atom positions does not have to produce an average of different placements. Individual samples can retain sharp local geometry while representing different configurations, although they are not automatically a physical conformational ensemble.

The shift also reflects a choice of representation. The authors found that simplifying AlphaFold 2’s structure module had only a modest effect on accuracy, while its residue frames and torsion angles were cumbersome for general molecular graphs. Raw atom coordinates offered a common representation across molecular types. I read diffusion as a useful way to learn and refine that representation; the AF3 paper does not report a direct comparison with an autoregressive coordinate decoder.

Boltz-2 and binding affinity

Boltz-1 provided an open model in the AlphaFold 3 family. Boltz-2 builds on recycled relationship representations and atom-coordinate diffusion, and adds binding-affinity prediction.

A predicted protein–ligand pose describes where a small molecule might sit. It does not establish how strongly the molecule binds. Boltz-2 adds a binding-likelihood output and an affinity-value output for the supported setting of small molecules binding to protein targets.

Representation recycling and coordinate denoising update separate states, followed by Boltz-2 affinity prediction.
Figure 7. Trunk recycling updates single and pair representations; diffusion updates atom coordinates. Affinity prediction addresses binding strength.

This helps compare candidate interactions rather than only predict their arrangements. The authors report favorable comparisons with free-energy perturbation on particular benchmarks, with performance varying by target. Those results do not establish a general replacement for physical calculations or experiments.

Thinking a bit more about recycling

Repeated learned updates predate AlphaFold 2. Universal Transformers shared attention and transition weights across depth in 2018. Later recurrent-depth language models use a shared transformer block to perform additional computation in hidden states.

What feels ahead of its time to me is that recycling scales inference FLOPs without scaling the parameter count. AlphaFold 2 reuses the same model across cycles, spending more compute to update its previous representations and geometry. Later looped transformers use a similar idea: run a shared block for more steps, increasing effective depth while keeping its parameters fixed. With R passes, the repeated computation costs roughly R times as many FLOPs as one pass. The parameters determine the update rule; the number of passes determines how much computation it gets to perform.

Shared computation

An ordinary stack can use a different function at each layer. A loop repeatedly applies the same function to a changing state:

hr+1 = Fθ(x, hr)

Here x is the input evidence, hᵣ is the working state, and θ denotes shared parameters. Weight sharing adds sequential computation without a proportional increase in parameter count, though each pass still costs time and compute. The shared block can learn an operation that is useful at multiple stages of solving the problem.

That is the connection to iterative algorithms. An intermediate result becomes context for the next application of the rule. Recent work demonstrates composition and depth extrapolation in controlled looped-transformer experiments. In AlphaFold, the intermediate result is a set of relationships and, in AlphaFold 2, a geometry hypothesis.

Refinement

A mathematical way to interpret useful refinement is a conditional fixed point: a state that the update leaves unchanged for a particular input. If updates locally reduce the error relative to that state, repeated passes can improve the hypothesis, with diminishing returns as the remaining error gets smaller.

Fixed evidence conditions a shared recurrent update, with an illustrative trajectory toward an input-dependent fixed point.
Figure 8. A shared update and an idealized refinement trajectory.

This is an explanatory model, not a proven property of AlphaFold. Weight sharing alone does not guarantee convergence, and a stable fixed point can still be incorrect. A model can also produce useful intermediate states without settling at the right answer.

Input conditioning

Refinement must retain the information that distinguishes one problem from another. Otherwise, a model could become increasingly stable while drifting toward an answer unrelated to its original input. The useful goal is to reduce answer errors while preserving distinctions between inputs.

AlphaFold keeps original evidence available during recycling. Some recurrent-depth language models similarly supply input embeddings to the recurrent block. This provides an anchor, but it does not guarantee that every pass uses it correctly.

Recent work finds that additional recurrence can enable deeper compositions, while excessive recurrence causes overthinking. Input injection did not resolve this in those experiments. More passes help only when the learned update remains useful; training depth and the point at which the output is read matter too.

Intermediate states

Recent work on LoopCD uses an earlier pass to guide decoding of the final pass. It contrasts predictions in logit space or extrapolates the final hidden state away from an earlier one, improving evaluated language-model benchmarks without retraining the recurrent block.

This suggests that changes between hypotheses can carry useful information alongside the latest state. They can also reflect noise or unstable dynamics. LoopCD supports using intermediate states for language-model readout; a similar approach to structural recycling would need separate experiments.

Diffusion performs a different repeated computation. Its updates follow a changing noise schedule, using the completed trunk as conditioning. Recycling revises the input representations; diffusion generates coordinates from them. Reusing parameters does not make those two loops the same operation.


Thank you to Thomas Bush, Bala Desinghu, and Bernardo Sabatini for the conversations and for helping revise this blog post.

References