Module reference¶
The full API surface, by category. Every function's docstring carries complexity, memory, and honest positioning; this page is the map.
Core¶
numba_utils.decorators¶
njit_fast—njit(cache=True, fastmath=True, nogil=True)for throughput kernelsnjit_parallel—njit(parallel=True, cache=True); read parallelism.md firstcached_njit— compile once, reuse across runs; see numba-cache.mdboundscheck— bounds checking withNUMBA_UTILS_DEV=1, plain njit in production
All accept bare and called forms; keyword overrides forward to njit
verbatim.
numba_utils.arrays¶
- Search over sorted arrays:
binary_search,lower_bound,upper_bound - Transforms:
fast_clip,normalize,cumulative_sum(all without=buffer reuse) - Windows:
rolling_sum,rolling_mean - Counting:
histogram(single pass, no edge array),bincount unique_sorted— dedup without the sort that dominatesnp.unique
numba_utils.algorithms¶
- Selection:
nth_element(in-place, C++ semantics),quickselect,fast_argpartition topk— heap path for small k, quickselect for large;argmax2(index AND value)- Sorts:
insertion_sort,partial_sort(in-place);counting_sort,radix_sort(new array; integer dtypes, honest loss vs NumPy's SIMD sort on full-range keys);stable_argsort(the stable-argsort spelling that works in nopython),lexsort(np.lexsortfor@njit— Numba doesn't implement it; takes a 2-D array, last row is the primary key) combination_table(n, k)— the C(n, k) index table; loop overtable.shape[0]instead of hardcoding combo counts (the evaluator bug class)disjoint_rank_aggregate/DisjointRankStructure— reach-weighted all-pairs comparison skipping pairs that share a key, EXACT via inclusion–exclusion over the 2^K−1 key subsets: O((2^K−1)·N log N) vs dense O(N·M). The value is the iterative shape:buildonce,evalper weight vector (CFR: 133x over dense with the build amortized) — the one-shot pays the full build (~80% of its time) and can lose to a dense pass; benchmark before adopting it for single evaluations. Domain envelope: K ≤ 12 distinct int64 keys per row AND C(V,K)·(K+1) < 2^62 over V distinct values (a 52-card deck fits at every K ≤ 12; the error message states the cap for your K). Certified against a dense reference with a drop-removal mutation that screams. Driven from Python, not njit-callable.
Performance¶
numba_utils.parallel¶
Complete parallel operations, not prange wrappers (docs,
design): parallel_sum, parallel_reduce
(per-index kernel decorator), parallel_histogram (bit-exact),
parallel_prefix_sum, parallel_topk. All fall back to serial below
SERIAL_THRESHOLD. Plus chunked_reduce — one per-chunk kernel,
serial and parallel drivers with bit-identical results: chunk
boundaries depend only on (n_items, n_chunks), never on thread
count; pair the chunk index with philox_uniform for runs that are
reproducible by construction.
numba_utils.profiling¶
benchmark— function mode excludes JIT compilation by default and times fast functions in auto-sized batches (TimingStats.inner), so per-call timer overhead never inflates the mean; block mode viawith benchmark():compare— two callables, same inputs, warmed up, samples INTERLEAVED with alternating order (drift lands on both sides): mean/median/variance + speedupwarmup(one call warms ONE signature),warmup_signatures(one call per dtype combination),compile_time,compile_stats
numba_utils.diagnostics¶
show(fn)— signatures, cache state, flags, compile timescheck(fn)— known-issue warnings with concrete recommendationsinspect(fn)— the underlying immutableFunctionReportshadowed()— modules loaded from a file the import path no longer resolves (a frozen snapshot, a stale build directory, a sibling checkout ahead onsys.path). Read-only, no dispatcher needed; scans first-party modules, or one module you name. The stale-artifact failure thatcache_locatorcovers for compiled binaries, one layer up at module resolution
Data structures¶
numba_utils.collections¶
jitclass-based, constructible and usable inside @njit
(design): Stack, FixedQueue, RingBuffer
(overwrite-oldest), PriorityQueue (binary min-heap), BitSet,
SparseSet (O(1) add/discard/contains/clear), ObjectPool (slot
allocator with double-release detection). Plus counter and
typed_defaultdict over typed dicts. Stack, FixedQueue,
RingBuffer and PriorityQueue are float64 by default; the
stack_type / fixed_queue_type / ring_buffer_type /
priority_queue_type factories return the same containers specialized
to any Numba scalar type, cached per type (stack_type(float64) is
Stack). Index-domain containers stay int64.
numba_utils.graph¶
Graph algorithms over CSR adjacency arrays (indptr, indices — the
scipy.sparse.csr_matrix layout; no Graph class, arrays are the
nopython-native currency): edges_to_csr (stable, returns an order
array to align per-edge payloads like weights), bfs (hop distances,
-1 unreachable), dfs_preorder (explicit stack, matches the recursive
order), topological_sort (Kahn, deterministic lowest-index-first,
raises on cycles), dijkstra (lazy-deletion binary heap, rejects
NaN/negative weights, inf = unreachable), and UnionFind (jitclass;
union by size + path compression, union returns whether a merge
happened). The CSR structure is validated up front (indptr monotonic
with the right endpoints) and indices entries are bounds-checked
during traversal — a malformed CSR raises instead of corrupting
memory.
numba_utils.stats¶
Numerically hard statistics — only functions whose naive versions are
wrong (design): logsumexp and softmax
(max-shifted; the direct formulas overflow exp beyond ~709; softmax
takes out=), weighted_quantile (inverted CDF — exact match
with np.quantile(..., weights=..., method="inverted_cdf"); rejects
NaN values and NaN/negative weights up front), and weighted_mc_mean
(uniform-subsample-then-weight, Philox-driven — the correct pattern
for the reach² bug that assert_no_reweight_bias guards against;
correct-pattern ≠ accurate at any n_sub: the ratio estimator's bias
grows with weight concentration, numbers in the docstring — certify at
your n_sub and weight profile).
numba_utils.random¶
Over Numba's nopython RNG, which is separate from NumPy's — seed it
with seed() (design): shuffle, permutation,
choice, reservoir_sampling (Algorithm R), sample_without_replacement
and the in-place partial_shuffle (partial Fisher–Yates, the
zero-allocation repeated-draw MC primitive), weighted_sampling, and
the Walker alias method as alias_setup / alias_draw /
alias_sample. Plus the stateless counter-based generator
(Philox4x64-10, bit-identical to np.random.Philox):
philox_uniform / philox_uniforms / philox_randint /
philox4x64 — pure functions of (key, counter), reproducible
regardless of threads, processes or call order — and the composed
variants philox_partial_shuffle / philox_sample_without_replacement
that drive the Fisher–Yates primitives from a Philox stream, four
swaps per block and in their own counter domain (so they consume
ceil(k / 4) counters and never collide with the three functions
above).
Developer tools¶
numba_utils.testing¶
assert_equivalent(reference, candidate, inputs)— per-case array copies, failing case named, empty generators failrandom_arrays— generated cases plus the edges that break kernelsassert_close,deterministic_rng(pins all three RNG worlds)- Stochastic asserts:
assert_reproducible(same seed → bit-identical) andassert_converges(different seeds → within N standard errors of the truth; the statistic is Student-t, real false-positive rates documented pern_runs). Both takepass_seed=Truefor counter-based (Philox) kernels, whose stream comes from an argument rather than global state.assert_within_seis the one-sample-set primitive underneath. - Certification:
mutation_screams(deliberately break the kernel, assert the check FAILS — a check that cannot fail certifies nothing; in-place kernels must return their buffer, and identical NaN/inf in both runs does not count as a scream) andassert_no_reweight_bias(screams on the reach² double-weighting bug; a pass requires the run to be CONCLUSIVE — an estimator too noisy to distinguish correct from double-weighted fails as inconclusive instead of certifying nothing)
Strategy: testing.md.
numba_utils.cache_locator¶
ContentHashLocator— stamps Numba's on-disk cache by SHA-256 of the source bytes instead of(mtime, size), closing the stale-binary window that mtime-preserving deployments (docker COPY,tar -x,rsync -a,cp -p) open. Opt-in via Numba'sNUMBA_CACHE_LOCATOR_CLASSEShook, set before the first import; not imported by the package itself. Full story: numba-cache.md.
Configuration¶
Global policy for cache / fastmath / parallel / nogil, from code
or environment. Overrides beat per-call arguments by design
(design).