Rectified flow superimposed visualization requires rectified flow data with at least 2 steps.
Flow-based generative models underlie many of the most capable modern generative systems (from image synthesis to video generation) and have evolved through a series of innovations, each addressing key limitations of the previous generation. A foundational entry point is normalizing flows, popularized by Rezende and Mohamed [?] , which transform a simple base distribution into a complex data distribution through a sequence of invertible, differentiable transformations. This invertibility enables exact likelihood computation via the change of variables formula, a significant advantage, but early normalizing flows required carefully constrained architectures to keep the Jacobian determinant tractable, limiting their expressiveness. Continuous normalizing flows (CNFs) [?] lifted these architectural constraints by taking the continuous-time limit: rather than a discrete chain of transformations, a CNF parameterizes a smooth vector field whose integral curves carry samples from source to target. This continuous formulation is far more flexible, but training required simulating the full ODE trajectory, which is computationally expensive and difficult to scale. Flow matching [?] resolved this bottleneck with a simulation-free training objective: a neural network is trained to regress directly onto conditional velocity fields, making CNF training practical at scale and allowing the use of arbitrary noise distributions.
Yet a new challenge emerges with flow matching: the trajectories learned by trained models tend to be curved, and accurately simulating these curved paths requires many costly neural network evaluations. For large models—often with billions of parameters—this incurs not just high computational cost but also high latency; in some cases it can take minutes to generate a single sample. This article traces the conceptual arc from normalizing flows through CNFs and flow matching, culminating in rectified flows [?] , a simple approach that straightens the trajectories of flow models to enable significantly faster sampling.
A major culprit behind the high cost incurred when sampling from flow models stems from the geometric properties of the learned flows. It can be challenging to reason about high-dimensional data, but fortunately for us, we can gain an intuition about many of the important geometric properties of flows by visualizing them in low-dimensions. In fact, we can use the exact same algorithms used to train large-scale models to train simple 2D flows on toy distributions and reproduce many phenomena of practical interest. In a related project we developed an interactive web app called Diffusion Explorer [?] that allows users to experiment with training and sampling from flow and diffusion models in 2D.
Sampling from a flow model involves simulating the trajectory of an abstract particle as it moves from random noise to real data by repeatedly querying a neural network to determine the particle's velocity at each point in time. When these trajectories are highly curved, accurately simulating them requires taking many small steps of our expensive neural network. Shown in Figure 2 above, we can see that a flow model trained to generate samples from a simple smiley face distribution produces trajectories that are curved. Understanding where this curvature comes from—and how to eliminate it—is the destination of the narrative arc this article traces. We begin with the foundations of normalizing flows and continuous normalizing flows, build up to flow matching as a scalable training framework, and then show how rectified flows [?] straighten these trajectories to enable fast, low-step sampling.
A normalizing flow [?] is a generative model that transforms a simple probability distribution, one that is typically easy to sample from, such as a multivariate Gaussian, into a complex target distribution through a sequence of invertible, differentiable transformations. The name reflects the idea that this sequence "flows" samples from a simple ("normal") distribution toward the complex data distribution. Concretely, given a base random variable drawn from a simple distribution , we apply a chain of transformations to obtain .
A central advantage of normalizing flows over many other generative models is their ability to compute exact likelihoods. Evaluating the density of a point under a simple distribution like a Gaussian is straightforward—we have a closed-form expression. But evaluating the density under a complex data distribution is not so easy. Normalizing flows solve this problem by providing a principled way to relate the density of a data point under the complex distribution to the density of its preimage under the simple base distribution, using the change of variables formula.
The Change of Variables Formula. The change of variables formula provides the mathematical link between the density of the base distribution and the density of the transformed distribution. For an invertible transformation that maps to , the density of is given by:
The term is the absolute value of the determinant of the Jacobian matrix of . It measures how much the transformation locally stretches or compresses volume. When the transformation expands a region of space, the density must decrease proportionally (and vice versa) so that the total probability mass is conserved. It is often convenient to work with the log form:
Jacobian Measures Local Volume Change. To build further intuition for the change of variables formula, we can visualize what the Jacobian determinant is actually measuring. The determinant of the Jacobian captures the local volume change induced by the transformation . Imagine a small region of space around a point —the Jacobian tells us how much this region is stretched, compressed, or rotated as it is mapped to the region around . If the determinant is greater than one, the transformation is locally expanding space (and the density decreases). If it is less than one, space is being compressed (and the density increases).
Composing Multiple Transformations. In practice, a single transformation is rarely expressive enough to bridge the gap between a simple base distribution and a complex data distribution. Instead, normalizing flows compose multiple transformations , each contributing a small change. The change of variables formula extends naturally to compositions—the log-likelihood under the full flow is the log-likelihood under the base distribution plus the sum of log-determinants at each stage:
Computing the Likelihood of Data. To compute the likelihood of an observed data point , we map it backward through the inverse transformations to recover its representation in the base space. At each step, we accumulate the log-determinant of the Jacobian of the inverse transformation. This gives us the log-likelihood of the data point:
Maximum Likelihood Training. With the ability to compute exact log-likelihoods, we can train normalizing flows by maximizing the log-likelihood of observed data. Given a dataset , we optimize the parameters of our flow transformations to maximize:
For each data point, we invert the flow, evaluate the base distribution density, and accumulate the log-determinant corrections. This provides an exact gradient signal for training, in contrast to models that rely on approximate inference.
While normalizing flows provide an elegant framework for exact likelihood computation, there is a major practical obstacle: computing the determinant of the Jacobian matrix is expensive. For a transformation in dimensions, computing the full Jacobian determinant requires operations in general:
This cubic cost is prohibitive for high-dimensional data like images, where can be in the tens of thousands or more. A substantial body of work has addressed this by restricting the form of the transformations to make the Jacobian determinant cheaper to compute—for example, planar flows, Real NVP, and Glow use architectures with triangular Jacobians, reducing the determinant to a product of diagonal entries. However, these restrictions limit the expressiveness of the flow. An alternative approach, which we discuss next, is to move to a continuous-time formulation that avoids the Jacobian determinant entirely.
Rather than composing a fixed number of discrete transformations, continuous normalizing flows (CNFs) (Chen et al., 2018) replace the sequence of layers with a single continuous-time ordinary differential equation (ODE) parameterized by a neural network. Instead of applying separate transformations, we define a smooth trajectory through the space of probability distributions.
A flow model learns to bridge a simple source probability distribution that is easy to draw samples from, like a multivariate Gaussian , to a complex data distribution by defining a continuous transformation between the two. We define a continuous sequence of probability distributions, called a probability path , that smoothly interpolates between our simple source distribution and our data distribution (see Figure 11). We index this path by a time variable , where corresponds to the source distribution and corresponds to the target distribution. By drawing samples from and transforming them over time we can produce samples distributed according to our data distribution .
Individual samples drawn from the source distribution have trajectories that trace a path from to as time progresses from to . These trajectories are the paths that individual "particles" take as the continuous normalizing flow transforms the source distribution into the target distribution. Understanding these trajectories is important because the geometry of these paths—how straight or curved they are—will turn out to have significant implications for the efficiency of sampling.
Rather than directly modeling the flow , continuous normalizing flows model a time-dependent velocity field that "generates" the flow. By taking this velocity field we can solve a set of ordinary differential equations (ODEs) to recover the flow, in a process called simulation. By starting from some initial point at time , we can trace the trajectory of this point over time according to the velocity field using the following ODEs
The solution to these ordinary differential equations involving is itself the flow . There are a variety of numerical methods for simulating these ODEs which approximate the continuous trajectory by taking a series of discrete steps. Perhaps the simplest such method is Euler's method, which approximates the trajectory of the flow by taking small linear steps in the direction of the velocity field at each time step .
A key advantage of the continuous-time formulation is that it enables more efficient likelihood computation. Recall that discrete normalizing flows require computing the full Jacobian determinant at each layer, which costs . In the continuous setting, the instantaneous change of variables formula replaces the expensive determinant with a much cheaper trace:
The trace of the Jacobian is only , a dramatic improvement over the cost of the full determinant. This makes CNFs practical for higher-dimensional data where discrete normalizing flows with unrestricted architectures would be computationally prohibitive.
Despite this efficiency gain, likelihood-based training of CNFs still has a significant drawback: it requires solving an ODE at every training step. To compute the log-likelihood of a data point, we must integrate the trace of the Jacobian along the entire trajectory from to , which requires solving the ODE for each training example. This simulation is computationally expensive and becomes a bottleneck during training.
Now that we have discussed the foundations of normalizing flows and continuous normalizing flows, we can discuss flow matching—a simulation-free training method that avoids the expensive ODE solving required by likelihood-based training. Please check out [?] for a more thorough introduction.
The motivation behind flow matching is to be able to learn our vector field without having to do expensive simulation, meaning without having to use Euler integration or some other technique to solve ODEs. Flow matching allows us to learn by solving a simple regression loss!
Flow matching can be broken down into two key steps:
We will focus on a specific choice of probability path called the linear path. The linear path can be defined through a simple linear interpolation between our source and target distributions:
In the examples I provide throughout this article, our source distribution is always a standard Gaussian distribution , and our target distribution is a complex 2D distribution representing a smiley face. However, in general, flow matching affords much more flexibility in the choice of probability paths and source distributions.
Now, the second step of flow matching is to "match" the true velocity field with an approximation , parameterized by a neural network, by optimizing a simple regression objective.
However, there is a catch: we do not have direct access to the true velocity field ! is difficult to directly construct in practice as it governs the transformations between two jointly distributed high dimensional distributions. So, how can we optimize this objective?
Luckily, we can create a related but much simpler objective by conditioning our velocity field on a particular instance from our target distribution . This yields the conditional velocity field .
Equipped with this conditional vector field, we can create a regression objective called conditional flow matching.
If we then plug in our specific conditional velocity field for our choice of a linear probability path, we get the remarkably simple training objective:
Incredibly, the conditional flow matching and the flow matching objectives have the same gradients , meaning we can optimize our tractable conditional flow matching objective and solve the flow matching problem. During training we simply need to draw pairs from our source and target distributions, interpolate between them to get , and then train our velocity field to predict the straight-line velocity .
A critical fact that is worth emphasizing, is that we are matching the conditional velocity which is conditioned on the target point with our learned velocity field which only "knows" about the current . If we were to condition our learned vector field on as well, then the problem would become trivial as the model could just predict some scaled version of . So, the model has to identify the likely destination using only the information about the location at time .
Stochastic interpolants [?] generalize the flow matching framework by adding controlled noise to the interpolation path. Rather than following a purely deterministic linear path between source and target, stochastic interpolants introduce a noise term where :
The noise schedule controls how much stochasticity is introduced at each time step. When for all , we recover the deterministic flow matching setting. When , the interpolant becomes stochastic, bridging the gap between deterministic flow models and stochastic diffusion models. This unifying perspective reveals that many seemingly different generative modeling approaches are special cases of a single framework.
Two Frameworks, One Idea. Remarkably, the flow matching framework [?] and the stochastic interpolants framework [?] were developed independently and in parallel, arriving at the same core insight: that one can train continuous normalizing flows by regressing onto conditional velocity fields without simulation. Despite different mathematical formulations and notation, both frameworks provide simulation-free training objectives that are equivalent under appropriate choices of interpolation schedules.
With the fundamentals of flow models and flow matching established, we can now investigate some of their idiosyncrasies—and how they come up in practice. We showed above that the trajectories produced by a flow model trained with flow matching are curved (see Figure 2). To further illustrate this point, if we superimpose the source and target distributions we can see that this curvature is even more extreme (see Figure 19).
Loading curved trajectory visualization...
An astute reader might recall that we trained our velocity field to match straight trajectories due to our choice of a linear path. So why does our model then learn curved trajectories, and why is this an issue? Answering the latter question—why curvature is a problem—is more straightforward: the answer is speed.
When drawing new samples from a flow model we perform numerical integration using the trained velocity field . At their core, numerical integration algorithms like Euler's method involve making finite steps in the direction of the velocity field: . We are making local linear approximations of the "true trajectories". The degree to which this approximation is accurate depends on how curved the trajectories are, and the size of steps we can take without deviating from the true trajectory, degrading sample quality.
The punch line: curvature is the enemy of speed. Highly curved trajectories are challenging to accurately simulate with a small number of steps. This means we need to make many calls to our large neural network representing our vector field in order to accurately approximate these trajectories, leading to high latency and computational cost. But why does our model learn these curved trajectories in the first place? The answer has to do with how our source and target random variables are jointly distributed, a concept called a coupling.
When training our velocity field with flow matching, we need to draw pairs from our source and target distributions and . Something that we glossed over a bit in the section about Flow Matching is how exactly we should draw these pairs. This is actually a crucial design choice, called a coupling, that has a significant impact on the geometry of the learned flow, and is the key culprit behind our curved trajectories.
A coupling is the joint distribution between our source and target random variables. This coupling dictates how our pairs used during training are distributed. The key requirement of a coupling is that the marginals are the source and target distributions .
The simplest form of coupling, and the one we investigate in this article, is an independent coupling (see Figure 21), where we independently draw and , and we have that . This allows us to trivially construct pairs during training, and is a natural choice in scenarios where we don't have any known structure associating pairs from our source and target distributions.
As mentioned above, our choice of independent coupling is the key culprit behind our curved trajectories. You can see in Figure 21 that the lines connecting independently drawn source and target points cross each other a lot. These intersections lead to curved trajectories because they introduce branches in our paths that our learned velocity field can not resolve.
An alternative to the independent coupling is an optimal transport coupling (see Figure 22), which connects source and target points in a way that minimizes the overall cost of transporting mass from the source to the target distribution. This coupling tends to produce fewer crossing paths, which leads to straighter trajectories. However, optimal transport couplings are more challenging to compute, especially in high dimensions, and so they are less commonly used in practice.
Our learned flow model is not capable of accurately modeling the crossing paths produced by our independent coupling; this incapability manifests itself in curved trajectories. More precisely, say two paths formed by the pairs and intersect at some point at time , or at least nearly intersect. This results in two distinct velocities and that our learned velocity field is supposed to match at the same location and time . This is not possible because our learned velocity field is only a function of the current location and time .
Our learned velocity field cannot accurately predict both desired velocities at this intersection point, and so it ends up predicting the average of these two velocities. This is also true more generally, whenever we have many paths intersecting in a small neighborhood. Our learned velocity field averages out the conflicting velocities by taking the conditional expectation of velocities passing through this point: . Finally, because the average velocity at these intersection points can change as we move through space, we develop curved trajectories. So, despite the fact that we train our flow model to match straight-line velocities, we end up with curved trajectories.
We have discussed why curved trajectories are difficult to simulate with a small number of steps, and now we also understand why a flow model learns curved trajectories when using an independent coupling. Now we ask the question: how can we learn straighter trajectories? A solution to this problem is exactly what Rectified Flows provide us with, and given all of the context above it is actually a startlingly simple solution laying in plain sight.
Rectified flows straighten out the trajectories of flows by replacing the naive independent coupling used in vanilla flow-matching training with one induced by the model itself. First, we train a model with flow matching using an independent coupling. Next, we generate new pairs by drawing and applying our learned flow model to get . This new coupling is then used to retrain a new flow model . By repeating this process multiple times we can progressively straighten out the trajectories of our flow model. The full procedure is outlined in the algorithm below.
We draw samples from our trained flow model by solving an ordinary differential equation of the form
This forms a deterministic flow, where it is guaranteed that trajectories are unique (under some mild regularity conditions). This uniqueness property is crucial to understanding why rectified flows work. The uniqueness of trajectories in deterministic flows means that two distinct trajectories cannot intersect at the same point in space and time . If this did happen, then the two trajectories would have to coincide for all times, contradicting the assumption that they are distinct. Deterministic flows therefore forbid crossing, branching, or merging of trajectories. The deterministic nature of these flows is inherited by the coupling induced by integrating the flow.
When we generate new pairs by flowing samples from our source distribution through our learned flow model, we are guaranteed to get a coupling where trajectories do not intersect. By retraining on this coupling, we are effectively removing the conflicting velocities at intersection points that caused curvature in the first place.
We can also compare the trajectories learned by a standard flow matching model versus a rectified flow model (see Figure 25 ). The rectified flow model learns significantly straighter trajectories, which are easier to simulate with fewer steps.
This difference in curvature has a direct impact on how many steps are needed during sampling. We can observe this effect by comparing how well Euler's method approximates the "ground truth" trajectory (using many steps) with varying numbers of integration steps (see Figure 26 ). Notice how the rectified flow model produces accurate approximations even with very few steps, while the flow matching model's curved trajectories lead to significant deviation from the true path.
Finally, we can compare the vector fields learned by a standard flow matching model versus a rectified flow model (see Figure 27 ). The rectified flow model learns vector field that is more consistent over time, meaning the model has lower curvature in its trajectories.
I'd like to acknowledge my friend Sebastián Gutiérrez Hernández for his valuable feedback on this project, particularly on the formal explanations presented in this article. I would also like to thank Benjamin Hoover, Polo Chau, and Vivek Anand for their feedback on the visualizations and writing.
If you found this explainer helpful, please consider citing it:
@article{helbling2026flowsurvey,
title = {A Visual Survey of Flow-Based Generative Models},
author = {Helbling, Alec},
year = {2026},
url = {https://alechelbling.com/qualifier-writeup}
}