Class anira::LatencyCalculator#

class LatencyCalculator#

Collaboration diagram for anira::LatencyCalculator:

digraph {
    graph [bgcolor="#00000000"]
    node [shape=rectangle style=filled fillcolor="#FFFFFF" font=Helvetica padding=2]
    edge [color="#1414CE"]
    "1" [label="anira::LatencyCalculator" tooltip="anira::LatencyCalculator" fillcolor="#BFBFBF"]
    "2" [label="anira::LatencyCalculator::Rational" tooltip="anira::LatencyCalculator::Rational"]
    "1" -> "2" [dir=forward tooltip="usage"]
}

Latency, inference-slot count and ring-buffer sizing of a session, in closed form.

Everything a SessionElement needs to size itself for a host configuration is a function of five numbers, all derived from the InferenceConfig and the HostConfig:

  • \(\rho = B / R\), the host block measured in hops: \(B\) is the host buffer size in samples of the reference stream and \(R\) that stream’s samples per inference (HostConfig::get_reference_size). \(\rho\) is stored as a reduced fraction \(p / q\) (see rationalize()).

  • \(\kappa = T \cdot f_s / (1000 \cdot R)\), the maximum inference time \(T\) (ms) measured in hop periods ( \(f_s\) is the host sample rate of the reference stream).

  • \(\beta\), the blocking ratio: the driving thread waits \(\beta\) host block periods for results inside every callback.

  • \(n\), the number of parallel processors of the session.

  • \(H\), the largest hop of any streamable tensor: the stream in which the host’s block size is finest, and the unit of the allow_smaller_buffers block grid.

The model is the scheduler’s actual worst case: every inference takes exactly \(T\), at most \(n\) run at once in submission order, an inference is submitted by the host callback that completes its hop, results are collected once per callback (after the blocking wait), and the host pushes and pops a sample only once a whole one has accumulated in its block (the documented fractional-block convention, so that per-stream block sizes may be fractional).

All of this is computed once, in the constructor; the getters are trivial. The class is public so that the formulas can be unit-tested against a brute-force simulation.

Buffer adaptation (Rath & Geier, LAC 2026)

Repackaging a stream from constant host blocks of \(b\) samples into blocks of \(P\) samples needs a delay of exactly \(\Delta = P - \gcd(b, P)\) samples (their corollary 3.8, replacing the PortAudio-style LCM loop). In hop units this is \(1 - 1/q\) for every stream at once, because \(\gcd(\rho P, P) = P / q\) whenever \(\rho P\) is an integer. When the host block may vary (allow_smaller_buffers) or is fractional on a stream, the worst case is \(P - 1\) samples (their section 5).

Inference queue

Let \(\tau = \kappa / \rho\) be the inference time in host blocks and \(d_m = \max(0, \lceil (m + 1)\tau - \beta \rceil)\) the number of callbacks after which the \((m+1)\)-th batch of \(n\) inferences submitted together is collected. The receive ring never runs dry iff its zero priming \(L\) (in hops) satisfies \(L \ge (k+1)\rho - C(k)\) for every callback \(k\), where \(C(k)\) counts the inferences collected by callback \(k\). With FIFO departures \(F_j = \max(a_j, F_{j-n}) + \tau\) and submissions \(a_j = \lceil (j+1)/\rho \rceil - 1\) this gives \(C(k) = \min_m [\lfloor \rho (k + 1 - d_m) \rfloor + m n]\), hence

\[ \Lambda = \max_k [(k+1)\rho - C(k)] = \frac{q-1}{q} + \max_{m \ge 0} [\rho\, d_m - m n], \]
the first term being the buffer adaptation above. The maximum over \(m\) exists iff \(\kappa < n\) (the pool keeps up with the stream); only \(m < \rho / (n - \kappa)\) can contribute, so the search is finite. The latency of output tensor \(i\) in its own samples is \(P_i \Lambda\) (see get_output_latencies()), rounded down because the host pops whole samples only.

Inference slots

By the same recursion the number of inferences submitted but not yet collected right after the submissions of a callback is at most

\[ S = \max_{m \ge 0} \left[ \lceil (d_m + 1)\rho \rceil - m n \right], \]
the steady-state slot count (get_num_structs()); the session allocates twice as many ThreadSafeStructs so that a wait-free reset, which strands the in-flight inferences in their slots until the workers finish, never starves the fresh schedule.

allow_smaller_buffers

The host may then use any block of \(j\) samples of the finest stream, \(1 \le j \le \lfloor \rho H \rfloor\), i.e. \(\rho' = j / H\) hops, with \(\tau' = \kappa / \rho'\). The worst case over that grid uses the flexible-host adaptation \((H - 1)/H\) and the maximum of \(\rho' d_m(\rho') - m n\) and of the slot count over \(j\). On every interval where \(\lceil (m+1)\tau' - \beta \rceil\) is a constant \(c\) both are increasing in \(j\), so the maximum is attained at the last grid point of the interval, \(j_c = \lceil (m+1)\kappa H / (c - 1 + \beta) \rceil - 1\), and the remaining intervals are bounded by \((m+1)\kappa\, c / (c - 1 + \beta)\). This replaces the former countdown over every block size.

Public Functions

LatencyCalculator(const InferenceConfig &inference_config, const HostConfig &host_config)#

Computes every quantity for one host configuration.

Parameters:
  • inference_config – Model configuration (hops, max inference time, blocking ratio, parallel processors)

  • host_config – Host configuration (buffer size and sample rate in reference-stream samples, allow_smaller_buffers, reference selection)

Throws:

std::invalid_argument – if the host config’s reference stream cannot be resolved

double get_latency_hops() const#

The worst-case latency in hops, \(\Lambda\) (fixed host block)

Returns:

Latency in hops, as a floating-point number

double get_latency_hops_smaller_buffers() const#

The worst-case latency in hops over the smaller-block grid.

Equal to get_latency_hops() unless HostConfig::m_allow_smaller_buffers is set.

Returns:

Latency in hops, as a floating-point number

std::vector<float> get_output_latencies() const#

Per-output float latency in samples of that output.

\(P_i \Lambda\) for every streamable output tensor \(i\), 0 for a non-streamable one; the smaller-block grid is included when the host config allows smaller buffers. Index-aligned with the output tensor list.

Returns:

Latency values in samples, one per output tensor

std::vector<unsigned int> get_synced_output_latencies() const#

Per-output integer latency, synchronized across outputs.

\(\lfloor P_i \Lambda \rfloor\) per streamable output; with more than one output tensor the latencies are then raised to a common whole number of hops (see sync_latencies()). Non-streamable outputs report 0. This is what SessionElement primes its receive rings with, before the internal model latency and any custom latency are applied.

Returns:

Latency values in samples, one per output tensor

size_t get_num_structs() const#

The steady-state number of inference slots, \(S\).

The maximum number of inferences submitted but not yet collected at any callback. SessionElement allocates twice this many ThreadSafeStructs: a wait-free reset leaves the in-flight inferences in their slots until the workers finish, while the fresh schedule needs \(S\) slots of its own.

Returns:

Maximum number of inferences in flight at any callback, \(S\)

std::vector<size_t> get_send_buffer_sizes() const#

Send ring sizes per input tensor.

One host block plus the largest leftover the adaptation can leave in the ring ( \(P - \gcd(\lceil b \rceil, P)\) for a constant integer block, \(P - 1\) for a fractional or variable one) plus the history a receptive-field model peeks at. 0 for a non-streamable input.

Returns:

Ring sizes in samples, one per input tensor

bool is_feasible() const#

Whether the inference pool can keep up with the stream.

False when \(\kappa \ge n\): every hop of audio brings \(\kappa\) hop periods of inference work for \(n\) processors, so the queue grows without bound and no finite latency covers it. The other getters then describe one host block processed by an idle pool, the best that can be said.

Returns:

True if the configuration is feasible

Rational get_block_hops() const#

The host block in hops, \(\rho = B/R\) as a reduced fraction.

double get_inference_hop_periods() const#

The inference time in hop periods, \(\kappa\).

Public Static Functions

static int64_t greatest_common_divisor(int64_t a, int64_t b)#

Greatest common divisor.

Parameters:
  • a – First non-negative integer

  • b – Second non-negative integer

Returns:

gcd(a, b), with gcd(0, b) = b

static int64_t least_common_multiple(int64_t a, int64_t b)#

Least common multiple, overflow-safe.

Divides before multiplying, \(a / \gcd(a, b) \cdot b\), so the intermediate never exceeds the result; the former a * b / gcd overflowed int once the product passed \(2^{31}\).

Parameters:
  • a – First non-negative integer

  • b – Second non-negative integer

Returns:

lcm(a, b), 0 if either is 0

static int64_t buffer_adaptation(int64_t host_block_size, int64_t stream_size)#

Minimum delay for repackaging constant host blocks into stream blocks.

Rath & Geier, corollary 3.8: \(\Delta = P - \gcd(b, P)\). Replaces the PortAudio-style loop over all multiples of \(b\) below \(\mathrm{lcm}(b, P)\).

Parameters:
  • host_block_size – Host block size in samples of the stream, \(b > 0\)

  • stream_size – Samples per inference of the stream, \(P > 0\)

Returns:

The delay in samples

static int64_t buffer_adaptation_flexible(int64_t stream_size)#

Minimum delay when the host block size varies or is fractional.

Rath & Geier, section 5: with flexible host blocks the last overlapping host block may start one sample before the stream block ends, so \(\Delta = P - 1\).

Parameters:

stream_size – Samples per inference of the stream, \(P > 0\)

Returns:

The delay in samples

static Rational rationalize(double value, int64_t max_denominator = 1 << 20)#

Best rational approximation of a non-negative value.

Continued-fraction convergents, stopping at the first within a relative tolerance of 1e-6 (the precision of the float host buffer size) or at the denominator bound.

Parameters:
  • value – The value to approximate, >= 0

  • max_denominator – Largest denominator to consider

Returns:

The approximation as a reduced fraction

static std::vector<unsigned int> sync_latencies(const std::vector<unsigned int> &latencies, const std::vector<size_t> &output_sizes)#

Synchronizes integer latencies across several output tensors.

With one output the value is kept. With several, every streamable output is raised to the same whole number of hops, \(\lceil \max_i L_i / P_i \rceil \cdot P_i\); non-streamable outputs stay at 0. Same rule as before this class existed.

Parameters:
  • latencies – Integer latencies in samples, one per output tensor

  • output_sizes – Postprocess output sizes (hops), one per output tensor

Returns:

Synchronized latencies in samples, one per output tensor

struct Rational#

A non-negative rational number.

Public Members

int64_t m_numerator = 0#

Numerator.

int64_t m_denominator = 1#

Denominator, always > 0.