July 20th 2026
RSS News

Structural Biology Has More to Teach AI Than Atomic Coordinates

The Reciprocal Space Station joins the Open Molecular Software Foundation to bring raw signals and frontier structural biology experiments to the next era of biomolecular AI.

Scroll

The first wave of biomolecular AI learned from atomic models and sequencing data. That was the right place to start: deposited structures have been the common language of structural biology to date, enabling decades of drug discovery, simulation and protein design. It was natural that structure prediction was the first big breakthrough in AI-powered bioscience. But experiments do not begin with the coordinates that make up a structure.

The data behind the structure — every diffraction pattern, scattering event, particle image — contains more information than can be captured in a single conformation. These data contain information about what molecules actually do: how they move, fluctuate, bind, and respond to perturbation. Coordinate models, however, need to enumerate all of that detail in an ever-growing list of atom positions, which is possible but already challenging for two or three discrete conformations. At the frontier of structural science are continuous motion and disorder — here, coordinate models become intractable: the number of parameters explodes beyond the power of the data to determine them, and the resulting models inhibit rather than aid scientific interpretation.

As a consequence, structural biologists seek the simplest, often single, conformation that explains as much of the data as possible, treating any variability — including functional variation — as noise. Tellingly, current structure predictors struggle with experimental heterogeneity in part because they learned from this compressed information. They trained on the world as it was deposited, not as it was measured.

This simplification causes us to lose information that holds the key to some of the most important open questions in molecular biology:

  • When is a mutation harmless, and when does it reshape a functional state?
  • Where are cryptic pockets, and how can we target them?
  • How do proteins transmit and transform information through their structures?
  • Can we understand the details of how water interacts with proteins — and use that to increase ligand potency?
  • How do molecular structure, dynamics, and function change with cellular context?

Some of the answers to these questions are hidden in the data we already have and collect today. Others will be the central pursuit of the experiments we invest in tomorrow.

To extract them, we need to go beyond the coordinates that we use to represent structure. Coordinate models are hand-fit to experimental data, built by humans, for humans. They have been incredibly successful: the Protein Data Bank (PDB) contains >250k structures that produced 9 Nobel prizes, guided >80% of approved drugs, and trained AlphaFold, hailed as the first scientific breakthrough powered by AI. But lists of coordinates also limit the ability to reason about essential but complex aspects of structure: allostery, dynamics, water networks, partially ordered and disordered regions — in short, heterogeneity.

We see a future in which AI models enable us to recapture this information, tackle the complexity of heterogeneity, and thereby fundamentally change how structure impacts biology and medicine. Our vision is to eventually let AI models interface directly with structural experiments, taking on the scale and complexity that current coordinate models can't handle. In doing so, we can extend what's scientifically possible — uncovering patterns and connections beyond what manual modeling allows — while freeing researchers to focus their reasoning on the questions that matter most.

To see this future more concretely, consider one of the most challenging and exciting tasks facing drug hunters today: designing a molecular glue degrader that creates a new protein-protein interface between two proteins. This task is essentially intractable for current rational, structure-based drug design. However, with the assistance of a biomolecular foundation model capable of abstractly reasoning about how small molecule binding induces structural changes, what drives protein-protein interactions, and the conformational preferences of any final complex, this task could become routine.

Surpassing the limits of coordinates to realize this vision will require a return to the measurements themselves. We need to build the AI-structure interface, the computational infrastructure that will make experimental data usable for inference of structural heterogeneity and enable training of foundation models based on distributions of structure.

RSS × OMSF

In pursuit of this future, we are excited to announce that the Reciprocal Space Station, or RSS, is joining the Open Molecular Software Foundation (OMSF) to tackle that goal together: building open, production-grade software that turns structural biology experiments into dynamics and function. OMSF's help with governance, administration, software engineering, and everything open source will enable RSS to expand our mission and tackle the frontier of dynamic structural software.

But what is a Reciprocal Space Station anyway? Our cheeky name riffs off reciprocal space, the mathematical representation where many structural biology experiments naturally live. Atomic-resolution structural biology relies on scattering: X-rays, electrons, or neutrons bounce off a sample, and a detector records the information imprinted on that radiation. Due to that radiation's wave-like nature, the process is naturally described in reciprocal space, the Fourier-space world where crystallography, cryo-EM, diffuse scattering, and small-angle methods meet. We are setting out to explore this data space from our home base, the Reciprocal Space Station.

RSS is part of the larger structural biology community, and we are committed to supporting and sustaining it and ongoing structural software development. Stay tuned, as more details about how RSS plans to operate and contribute will follow in a future post!

Building a path from measurements to function

The bottleneck RSS is tackling is not just one missing algorithm, but a stack of connected problems. Experimental data needs a translation layer: from facility-specific files and scientific conventions to context-aware arrays and tensors. Forward models, which map conformations back to experimental observables, need to be accurate, differentiable, and fast. Finally, we need the ability to go beyond single structures to conformational distributions that change depending on environment, composition, and time.

Our projects target that stack directly: reciprocalspaceship for programmable crystallographic data; Careless, Meteor, and Laue-DIALS for difficult diffraction and time-resolved regimes; SFCalculator for differentiable forward modeling; and ROCKET and EmbedOpt for observable-guided structure and ensemble inference.

Together, these efforts define the RSS roadmap: make structural biology data accessible and interoperable, implement the underlying physics as differentiable models, and train the next generation of biomolecular AI on data, not structures.

Where we are going

We've already started down this path, with some exciting progress.

Ensemble fitting from raw observables

In dynamic regions, building a structural model by hand is often impossible. To address this challenge, we are developing engines that automatically fit distributions of conformations directly to experimental data. This will enable us to capture allostery, where changes on one side of a protein propagate to another, spatially removed site. Allostery opens new opportunities in drug discovery and is hard to capture with a single structure alone. We are therefore extending our machinery to ensembles — stay tuned for an upcoming post!

Data velocity for modern instruments

Once interpretation no longer requires hours of manual model building per dataset, scale becomes the opportunity. We are partnering with efforts such as OpenBind and OpenADMET to help turn modern high-throughput structural experiments into usable data streams for cofolding, binding, and ADMET modeling.

From guiding existing models to training new ones

Alongside sister projects such as OpenFold, we aim to close the experiment-model loop: measurements guide models, while models help design and interpret experiments. We plan to build predictors that use experimental data directly, not only to validate a final structure, but to guide structure determination, ensemble fitting, and model training. We believe these will enable the next big breakthroughs in biomolecular AI: cofolding, protein-protein interface prediction, and the design of protein dynamics.

Why OMSF

To realize our mission, scientific excellence is necessary but not sufficient. Automation, performance, usability, documentation, testing, education, and long-term stewardship are first-order concerns of RSS, alongside scientific insight. It will take collaboration across structural biology, computation, simulation, drug discovery, and experimental facilities, and code that's not just impactful, but available, reliable, extensible, and community-driven.

That's why RSS is joining the Open Molecular Software Foundation: a home for open molecular software projects that support entire fields. Alongside consortia like OpenFold, OpenForceField, OpenADMET, and OpenBind, RSS will help build the open infrastructure the next generation of structural biology needs.

Join us

At its heart, RSS is an open community of reciprocal astronauts: structural biologists, software developers, experimentalists, method builders, and partners who believe experiments have more to teach us than a single, static coordinate model.

If you are excited about open software, experimental data, molecular dynamics, and the future of structural biology, reach out at crew@rs-station.org, find us on GitHub, or join our community forum.

Structural biology has more to teach AI than coordinates. Let's build the tools to let it speak.

⬅️ Back to blog posts