Latency#

Overview#

Latency is a critical factor in real-time audio processing applications. anira computes the latency of every session in closed form, together with the number of inference slots the session allocates and the sizes of its ring buffers. The calculation lives in anira::LatencyCalculator; anira::SessionElement applies its result in prepare().

The model behind the formulas is the scheduler’s actual worst case:

  • every inference takes exactly max_inference_time,

  • at most num_parallel_processors inferences run at once, in submission order,

  • an inference is submitted by the host callback that completes its hop of input,

  • results are collected once per callback, after the blocking wait (blocking_ratio),

  • the host pushes and pops a sample only once a whole one has accumulated in its block, so that a block may be a fractional number of samples on a stream (see the note below).

Quantities#

For a session with reference stream size \(R\) (see the usage guide, section 4.1), host buffer size \(B\) and host sample rate \(f_s\), both stated in samples of the reference stream, the calculator derives:

Symbol

Value

Meaning

\(\rho = p / q\)

\(B / R\)

the host block measured in hops, as a reduced fraction

\(\kappa\)

\(T f_s / (1000 R)\)

the maximum inference time \(T\) (ms) measured in hop periods

\(\tau\)

\(\kappa / \rho\)

the inference time measured in host blocks

\(\beta\)

blocking_ratio

the fraction of a host block the driving thread waits for results

\(n\)

num_parallel_processors

the parallelism of the session

\(H\)

largest streamable hop

the stream on which the host block is finest; the unit of the allow_smaller_buffers grid

\(d_m\)

\(\max(0, \lceil (m+1)\tau - \beta \rceil)\)

callbacks after submission at which the \((m+1)\)-th batch of \(n\) inferences submitted together is collected

Buffer adaptation#

Repackaging a stream from host blocks of \(b\) samples into model blocks of \(P\) samples requires a delay. Rath and Geier [RG26] prove that the minimum delay for constant block sizes is

\[\Delta = P - \gcd(b, P),\]

which replaces the PortAudio-style loop over every multiple of \(b\) below \(\mathrm{lcm}(b, P)\) that anira used before. In hop units the adaptation is \(1 - 1/q\) for every stream at once, because \(\gcd(\rho P, P) = P / q\) whenever the block \(\rho P\) is a whole number of samples. If the host block may vary in size (allow_smaller_buffers) or is a fractional number of samples on a stream, the worst case is \(P - 1\) samples [RG26], section 5.

Inference-caused latency#

The receive ring of an output never runs dry if and only if its zero priming \(L\) (in hops) covers, at every callback \(k\), the samples the host has popped minus the results that have been collected:

\[L \;\ge\; (k+1)\rho - C(k) \quad \text{for all } k,\]

where \(C(k)\) is the number of inferences collected by callback \(k\). Inference \(j\) is submitted at callback \(a_j = \lceil (j+1)/\rho \rceil - 1\) (the last host block overlapping model block \(j\), the observation at the heart of [RG26]), and with \(n\) processors in submission order it finishes at \(F_j = \max(a_j, F_{j-n}) + \tau\). Unrolling the recursion gives

\[C(k) = \min_{m \ge 0} \left[ \lfloor \rho\,(k + 1 - d_m) \rfloor + m n \right],\]

and since \((k+1)\rho - \lfloor \rho (k+1-d_m) \rfloor = \rho\, d_m + \mathrm{frac}(\rho (k + 1 - d_m))\), whose largest value over \(k\) is \(\rho\, d_m + (q-1)/q\), the latency in hops is

\[\Lambda \;=\; \frac{q - 1}{q} \;+\; \max_{m \ge 0} \left[ \rho\, d_m - m n \right].\]

The first term is the buffer adaptation, the second the inference queue. The maximum exists if and only if \(\kappa < n\): every hop of audio brings \(\kappa\) hop periods of inference work for \(n\) processors, and beyond that the queue grows without bound (anira::LatencyCalculator::is_feasible() is false, prepare() logs a warning, and the values describe one host block processed by an idle pool). Only batches with \(m < \rho / (n - \kappa)\) can contribute, so the search is finite; in the common case \(\rho \lceil \tau \rceil \le n\) the maximum is the first term, \(\rho \lceil \tau - \beta \rceil\).

The latency of output tensor \(i\) in its own samples is \(P_i \Lambda\). It is the same number of hops for every output, and it is rounded down because the host pops whole samples only (anira::LatencyCalculator::get_output_latencies() returns the unrounded value).

Note

When the host buffer size is a fractional (floating-point) value on a stream, the host and that stream exchange samples at a non-integer ratio. anira assumes the worst case: a sample is pushed to the anira::InferenceHandler only when the host buffer accumulates a full sample, and popped under the same rule. For example, if the host buffer size is 0.25 samples of a control-rate stream, the anira::InferenceHandler receives one sample every four host buffer cycles, and the latency is calculated as if the sample is delivered during the fourth cycle. If your system always sends the sample at the first host buffer cycle, a lower latency is possible; consider configuring anira::InferenceHandler with a custom latency value.

Inference slots#

The same recursion bounds the number of inferences that are submitted but not yet collected right after the submissions of a callback, which is the number of thread-safe structures a session allocates:

\[S \;=\; \max_{m \ge 0} \left[ \lceil (d_m + 1)\rho \rceil - m n \right].\]

With a blocking ratio that covers the inference time this is 1: the result is collected in the callback that submitted it.

The session allocates twice this many structures. anira::InferenceHandler::reset() is wait-free: it does not wait for in-flight inferences but marks them stale, and each keeps its structure until its worker publishes completion. Meanwhile the fresh schedule needs \(S\) structures of its own, so one pool drains while the other serves and a reset never drops a hop.

Adaptive buffer handling#

For hosts that may use smaller buffers than the configured maximum (allow_smaller_buffers), the host block may be any whole number \(j\) of samples of the finest stream, \(1 \le j \le \lfloor \rho H \rfloor\), i.e. \(\rho' = j / H\) hops with \(\tau' = \kappa / \rho'\). The calculator takes the worst case over that grid: the flexible-host adaptation \((H-1)/H\) plus the maximum over \(j\) of the inference term and of the slot count. Both are increasing in \(j\) on every interval where \(\lceil (m+1)\tau' - \beta \rceil\) is a constant \(c\), so the maximum is attained at the last grid point of each interval, \(j_c = \lceil (m+1)\kappa H / (c - 1 + \beta) \rceil - 1\), and the remaining intervals are bounded by \((m+1)\kappa\, c / (c - 1 + \beta)\). Only these breakpoints are evaluated; the former countdown over every block size, which stalled for blocks above \(2^{24}\) samples, is gone.

Note that a smaller host block can require more latency and more slots than the maximum block: a block slightly shorter than one inference time means that results arrive two callbacks after their submission instead of one.

Internal model latency#

Additional latency inherent to the model itself, such as look-ahead requirements or internal buffering (internal_model_latency), is added to every streamable output after the calculation above.

Latency synchronization#

When several output tensors are present, the integer latencies are raised to a common whole number of hops, \(\lceil \max_i L_i / P_i \rceil \cdot P_i\), so that the outputs stay coherent.

The latency vector returned by anira::InferenceHandler::get_latency_vector() is index-aligned with the output tensor list. Non-streamable outputs (postprocess_output_size == 0) carry no stream latency and always report 0.

Ring buffer sizes#

A send ring holds one host block plus the largest leftover the adaptation can leave behind, \(\lceil b \rceil + P - \gcd(\lceil b \rceil, P)\) for a constant whole-sample block and \(\lceil b \rceil + P - 1\) for a fractional or variable one, plus the history a receptive-field model peeks at. A receive ring holds one result per inference slot plus the latency priming, \(S P_i + L_i\).

These calculations count buffer sizes in samples of the reference stream selected by the anira::HostConfig — a streamable input for effects and analysers, the streamable output for a generator model with no streamable input (see the usage guide, section 4.1).

Output Behavior#

The final latency value represents the total delay (in samples) between when input data enters the system and when the processed output data becomes available.

Important

Before the first valid output is produced, the anira::InferenceHandler::process() and anira::InferenceHandler::pop_data() methods will return zeroed data. This ensures real-time audio processing without introducing unexpected delays or artifacts in the output signal.

[RG26] (1,2,3)

Matthias Rath and Matthias Geier, “Minimum required delay for realtime block size adaptation in digital audio signal processing”, Proceedings of the 20th Linux Audio Conference (LAC-26), Maynooth, 2026. hal-05697688