Neural Network Primitives (BNNS)
AppleAccelerate wraps the current, non-deprecated slice of Apple's BNNS (Basic Neural Network Subroutines) library — 61 of the ~136 BNNS* C entry points. The bulk of the remainder are APIs Apple deprecated in macOS 15 / iOS 18: the classic filter/layer construction API and the deprecated classic + DirectApply tensor kernels (BNNSMatMul, BNNSTile, BNNSGather/BNNSScatter, the clip / norm family, BNNSOptimizerStep, …). Those are intentionally not wrapped — target the BNNS Graph API instead, which this page covers end to end: compile a Core ML model, inspect its arguments, and run inference on Julia arrays. Numerically verified helpers (transpose, copy, reductions, top-k, random generation, nearest neighbors, graph execution) are cross-checked against plain-Julia references; the remaining thin wrappers expose the rest of the current surface with exact FFI signatures for callers who need them.
Element types
Float32 works everywhere. Float16 and the integer / Bool types are accepted wherever the underlying kernel was verified at runtime to compute correctly — BNNS returns a success status with wrong values for some unsupported types, so each wrapper restricts its element types by dispatch:
| Function | Element types |
|---|---|
bnns_transpose, same-type bnns_copy! | Float16, Float32, Int8–Int64, UInt8–UInt64, Bool |
converting bnns_copy! | Float16 ↔ Float32, Int8/Int16/Int32/UInt8 → Float32, Float32 → Int32 (anything else throws) |
bnns_reduce | Float16, Float32; Int32 for the integer-exact reductions |
bnns_topk | Float16, Float32, Int8, Int16, Int32, UInt8, UInt16 |
bnns_in_topk | Float16, Float32 |
| uniform / normal / categorical random fills | Float16, Float32 |
bnns_random_fill_uniform_int! | Int8–Int64, UInt8–UInt64 |
| graph inputs / outputs | whatever the model declares: Float16, Float32, integers, Bool |
These functions are not exported. Access them via the AppleAccelerate. prefix (e.g. AppleAccelerate.bnns_reduce).
Apple deprecated the classic BNNS filter/layer API and much of the classic tensor/DirectApply surface (macOS 15 / iOS 18) in favour of the newer BNNS Graph API (BNNSGraph). This package does not wrap any of those deprecated entry points; use the Graph API for that functionality. The "What's left to the raw layer" section below lists the full excluded set.
Descriptors
BNNSArray builds a GC-safe BNNSNDArrayDescriptor view of a dense, contiguous Julia array. Internally the N-D op wrappers map a column-major Julia Array onto a BNNSDataLayout{N}DLastMajor descriptor with explicit strides, so BNNS axis k corresponds to Julia dimension k+1 (axis 0 is the contiguous/fastest axis).
AppleAccelerate.BNNSArray — Type
BNNSArray(A::AbstractArray)A GC-safe BNNSNDArrayDescriptor view of a Julia array A, suitable for passing to BNNS routines via Ref. The wrapper keeps a reference to the backing array so it is not collected while the descriptor is alive; pass the underlying descriptor with Base.cconvert/Ref only inside a GC.@preserve block guarding A.
BNNS descriptors are layout aware. Julia stores arrays in column-major order, so this constructor reports the array using a column-major-friendly BNNS layout:
- 1D
Vector->BNNSDataLayoutVector - 2D
Matrix->BNNSDataLayoutColumnMajorMatrix(BNNSsize = (rows, cols)).
Only Float16, Float32 and Int32 dense, contiguous arrays are supported here; other element types or strided/transposed arrays should use the raw LibAccelerate layer directly.
Tensor manipulation
Stateless tensor ops that remain current, cross-validated against permutedims and plain copies.
| Function | Meaning |
|---|---|
bnns_transpose | swap two axes |
bnns_copy! | copy, optionally converting the element type |
julia> M = Float32[1 2 3; 4 5 6];
julia> AppleAccelerate.bnns_transpose(M, 1, 2)
3×2 Matrix{Float32}:
1.0 4.0
2.0 5.0
3.0 6.0
julia> AppleAccelerate.bnns_copy!(zeros(Float32, 2, 3), M) == M
true
julia> AppleAccelerate.bnns_copy!(zeros(Float16, 2, 3), M) # converting copy: Float32 -> Float16
2×3 Matrix{Float16}:
1.0 2.0 3.0
4.0 5.0 6.0AppleAccelerate.bnns_transpose — Function
bnns_transpose(A::Array, dim0, dim1) -> ArraySwap Julia dimensions dim0 and dim1 of A (1-based) via BNNSTranspose, equivalent to a permutedims that exchanges those two axes. Works for every element type BNNS can describe: Float16, Float32, Int8–Int64, UInt8–UInt64 and Bool.
AppleAccelerate.bnns_copy! — Function
bnns_copy!(dest::Array, src::Array) -> destCopy src into the equally-sized dest via BNNSCopy. For equal element types this is a plain element copy.
dest and src may have different element types, in which case BNNS converts: Float16 ↔ Float32 (the usual way to move data in and out of a half-precision graph), Int8/Int16/Int32/UInt8 → Float32, and Float32 → Int32. Any other pair throws an ArgumentError without calling BNNS.
Reductions
AppleAccelerate.bnns_reduce — Function
bnns_reduce(func::Symbol, input::Array; dim=1) -> ArrayReduce input along Julia dimension dim with func (:sum, :mean, :max, :min, :sumsquare, :l1, :l2, :product, :logsumexp) via BNNSDirectApplyReduction. The reduced axis collapses to length 1 and the result has the element type of input.
Supported element types: Float32, Float16 (computed in half precision, so sums saturate at floatmax(Float16) = 65504) and Int32. For Int32 only the reductions that are exact in integers are offered (:sum, :max, :min, :sumsquare, :l1, :product); :mean, :l2 and :logsumexp throw.
DirectApply kernels
Fused kernels that run without an explicit filter handle.
AppleAccelerate.bnns_topk — Function
bnns_topk(input::Array, K; dim=1) -> (values, indices)Top-K values and their 0-based indices along Julia dimension dim via BNNSDirectApplyTopK. values has the element type of input, indices is Int32. Comparable to sort-based partialsortperm per slice.
Supported element types: Float32, Float16, Int8, Int16, Int32, UInt8, UInt16 (BNNS rejects the wider integer types).
AppleAccelerate.bnns_in_topk — Function
bnns_in_topk(input::Array, targets::Array{Int32}, K; dim=1) -> Array{Bool}For each batch column, test whether the targets class index is among the top-K scores of input along Julia dimension dim (BNNSDirectApplyInTopK). input may be Float32 or Float16.
Utility queries
AppleAccelerate.bnns_layout_rank — Function
bnns_layout_rank(layout::BNNSDataLayout) -> IntRank (number of dimensions) encoded by a BNNSDataLayout constant, via BNNSDataLayoutGetRank.
AppleAccelerate.bnns_data_size — Function
bnns_data_size(A::Array) -> IntNumber of bytes of tensor data described by A (BNNSNDArrayGetDataSize).
AppleAccelerate.bnns_tensor_allocation_size — Function
bnns_tensor_allocation_size(A::Array) -> IntBytes required to allocate a BNNSTensor describing A (BNNSTensorGetAllocationSize). Uses the modern BNNSTensor struct (rank + shape/stride), distinct from the legacy BNNSNDArrayDescriptor.
Random number generation
BNNSRandomGenerator is an AES-CTR generator with an optional seed; the fill functions populate arrays in place and the state can be snapshot and restored for reproducibility.
AppleAccelerate.BNNSRandomGenerator — Type
BNNSRandomGenerator([seed]) -> BNNSRandomGeneratorA BNNS random number generator handle (AES-CTR method). Construct with an optional 64-bit seed for reproducibility (BNNSCreateRandomGeneratorWithSeed, or BNNSCreateRandomGenerator when omitted). The handle is destroyed automatically by a finalizer (BNNSDestroyRandomGenerator).
Use with bnns_random_fill_uniform!, bnns_random_fill_normal!, bnns_random_fill_uniform_int!, bnns_random_fill_categorical! and the bnns_random_state/bnns_random_state! round-trip.
AppleAccelerate.bnns_random_fill_uniform! — Function
bnns_random_fill_uniform!(g::BNNSRandomGenerator, A::Array, lo=0f0, hi=1f0) -> AFill A (Float32 or Float16) with i.i.d. uniform samples on [lo, hi) (BNNSRandomFillUniformFloat). For Float16 the samples are rounded to half precision, so a value can round up to exactly hi.
AppleAccelerate.bnns_random_fill_uniform_int! — Function
bnns_random_fill_uniform_int!(g::BNNSRandomGenerator, A::Array{<:Integer}, lo, hi) -> AFill integer array A with i.i.d. uniform samples on the half-open range [lo, hi) (BNNSRandomFillUniformInt). A may be Int8, Int16, Int32, Int64, UInt8, UInt16, UInt32 or UInt64; the range must fit the element type.
AppleAccelerate.bnns_random_fill_normal! — Function
bnns_random_fill_normal!(g::BNNSRandomGenerator, A::Array, mean=0f0, stddev=1f0) -> AFill A (Float32 or Float16) with i.i.d. Gaussian samples (BNNSRandomFillNormalFloat).
AppleAccelerate.bnns_random_fill_categorical! — Function
bnns_random_fill_categorical!(g::BNNSRandomGenerator, out::Array{T}, probs::Array{T}; log_probs=false) -> outDraw categorical samples (0-based category indices, stored as floating point) into out using per-category weights probs (BNNSRandomFillCategoricalFloat). Pass log_probs=true if probs holds log probabilities. T is Float32 or Float16; out and probs must share it (BNNS silently mis-samples mixed precisions).
AppleAccelerate.bnns_random_state — Function
bnns_random_state(g::BNNSRandomGenerator) -> Vector{UInt8}Snapshot the generator's internal state (BNNSRandomGeneratorStateSize + BNNSRandomGeneratorGetState). Restore it with bnns_random_state!.
AppleAccelerate.bnns_random_state! — Function
bnns_random_state!(g::BNNSRandomGenerator, state::Vector{UInt8}) -> gRestore a generator state captured by bnns_random_state (BNNSRandomGeneratorSetState).
Nearest neighbors
AppleAccelerate.BNNSNearestNeighbors — Type
BNNSNearestNeighbors(max_samples, n_features, n_neighbors; T=Float32) -> BNNSNearestNeighborsA brute-force k-nearest-neighbours index (BNNSCreateNearestNeighbors) holding up to max_samples reference points of dimension n_features, answering n_neighbors-NN queries. Destroyed automatically (BNNSDestroyNearestNeighbors).
Add reference points with bnns_knn_load! and query with bnns_knn_query.
AppleAccelerate.bnns_knn_load! — Function
bnns_knn_load!(knn::BNNSNearestNeighbors, data::Matrix{Float32}) -> IntAppend reference samples to the index (BNNSNearestNeighborsLoad). data is n_features × n_new_samples (each column is one sample, matching BNNS's feature-major layout). Returns the number of samples loaded.
AppleAccelerate.bnns_knn_query — Function
bnns_knn_query(knn::BNNSNearestNeighbors, sample_number) -> (indices, distances)Return the n_neighbors nearest reference points to the (0-based) loaded sample sample_number (BNNSNearestNeighborsGetInfo): their 0-based indices (Vector{Int32}) and Float32 distances.
BNNS Graph API
The modern, non-deprecated pipeline (macOS 15+): compile a Core ML model into a BNNSGraph, make an executable BNNSGraphContext, look at what it expects with bnns_graph_arguments, and run it on Julia arrays with bnns_graph_run / bnns_graph_run!.
The input is a compiled Core ML model — the .mlmodelc directory that Xcode or xcrun coremlcompiler compile model.mlpackage out/ produces (ML Program models only). There is no in-memory graph builder in this API. A .mlmodelc is a directory holding a textual MIL program, model.mil, plus a weights blob; the example below writes a tiny one by hand so that it is self-contained.
modeldir = joinpath(mktempdir(), "dense.mlmodelc"); mkpath(modeldir)
write(joinpath(modeldir, "model.mil"), """
program(1.3)
[buildInfo = dict<string, string>({{"coremlc-component-MIL", "handwritten"}})]
{
func main<ios16>(tensor<fp32, [2, 3]> x) {
tensor<fp32, [3]> b = const()[name = string("b"), val = tensor<fp32, [3]>([1.0, -2.0, 0.5])];
tensor<fp32, [2, 3]> s = add(x = x, y = b)[name = string("s")];
tensor<fp32, [2, 3]> z = relu(x = s)[name = string("z")];
} -> (z);
}
""")
graph = AppleAccelerate.BNNSGraph(modeldir)
ctx = AppleAccelerate.BNNSGraphContext(graph)
AppleAccelerate.bnns_graph_arguments(ctx)2-element Vector{AppleAccelerate.BNNSGraphArgument}:
BNNSGraphArgument("z", :out, Float32, shape=(2, 3), size=(3, 2))
BNNSGraphArgument("x", :in, Float32, shape=(2, 3), size=(3, 2))Memory layout
MIL tensors are row-major, Julia arrays are column-major, so a model tensor of shape [2, 3] has the memory layout of a Julia array of size (3, 2). The graph functions pass Julia's memory to BNNS untouched, which means every argument is a Julia array whose size is the reverse of the model's shape — that is the size field reported above. For an image model, [N, C, H, W] is a Julia (W, H, C, N) array. BNNS graphs ignore custom strides, so this is the only zero-copy mapping.
X = Float32[1 2 3; -4 0 6] # in the model's [2, 3] index order
out = AppleAccelerate.bnns_graph_run(ctx, "x" => permutedims(X)) # (3, 2): reversed dims
@assert permutedims(out["z"]) == max.(X .+ Float32[1 -2 0.5], 0)
# mil_order=true does that permutedims for you, on the way in and on the way out
out = AppleAccelerate.bnns_graph_run(ctx, "x" => X; mil_order = true)
@assert out["z"] == max.(X .+ Float32[1 -2 0.5], 0)For repeated inference preallocate the outputs and a page-aligned workspace, and nothing is allocated per call on the BNNS side:
Z = zeros(Float32, 3, 2)
ws = AppleAccelerate.bnns_graph_workspace(ctx)
AppleAccelerate.bnns_graph_run!(ctx, "z" => Z, "x" => permutedims(X); workspace = ws)
@assert permutedims(Z) == out["z"]A half-precision model takes and returns Float16 arrays; convert with bnns_copy! (or plain Float16.(x)). Models with a dynamic batch dimension are bound with bnns_graph_context_set_batch_size!, more general dynamic shapes with bnns_graph_context_set_dynamic_shapes!. A context carries mutable state and must be used by one thread at a time; make one context per task.
AppleAccelerate.BNNSGraphCompileOptions — Type
BNNSGraphCompileOptions(; single_thread=nothing, generate_debug_info=nothing,
optimization=nothing, log_mask=nothing,
output_path=nothing, output_fd=nothing) -> BNNSGraphCompileOptionsOptions controlling BNNSGraphCompileFromFile, backed by BNNSGraphCompileOptionsMakeDefault and destroyed by a finalizer (BNNSGraphCompileOptionsDestroy). Any keyword left nothing keeps the BNNS default. optimization is :performance or :ir_size. Individual fields can also be read/written with the accessor functions below.
AppleAccelerate.BNNSGraph — Type
BNNSGraph(filename; func=nothing, options=BNNSGraphCompileOptions()) -> BNNSGraphCompile the compiled Core ML model (.mlmodelc directory, or the model.mil inside it) at filename — optionally only the named func inside it — into an executable graph via BNNSGraphCompileFromFile. Throws if BNNS cannot compile the model (unsupported op, malformed program, missing file). The returned handle feeds BNNSGraphContext and the graph-introspection helpers, and the compiled graph's memory is released by a finalizer once the graph and every context made from it are unreachable.
.mlmodelc is what Xcode / xcrun coremlcompiler compile produce from an .mlpackage; only ML Program models (MIL), not the older NeuralNetwork format, are accepted by BNNS.
AppleAccelerate.BNNSGraphContext — Type
BNNSGraphContext(g::BNNSGraph) -> BNNSGraphContextAn executable context for a compiled BNNSGraph (BNNSGraphContextMake), destroyed by a finalizer (BNNSGraphContextDestroy). It holds the mutable execution state (dynamic shapes, streaming state), keeps its graph alive, and must be used by one thread at a time; make one context per task for concurrent inference. Feed it to bnns_graph_run / bnns_graph_run!, or to the low-level bnns_graph_execute!.
AppleAccelerate.BNNSGraphArgument — Type
BNNSGraphArgumentDescription of one argument of a graph function, as returned by bnns_graph_arguments:
name::Stringintent::Symbol—:in,:outor:inouteltype::DataType— Julia element type (Float16,Float32,Int32,Bool, …)shape::Dims— the shape as written in the model (MIL / row-major order)size::Dims—sizeof the JuliaArrayto pass for it:reverse(shape)
A dimension of 0 (in either tuple) is dynamic and not bound yet.
AppleAccelerate.bnns_graph_arguments — Function
bnns_graph_arguments(c::BNNSGraphContext; func=nothing) -> Vector{BNNSGraphArgument}
bnns_graph_arguments(g::BNNSGraph; func=nothing)Names, intents, element types and shapes of every argument of func, in execute order (outputs first). Given a context, shapes reflect any batch size / dynamic shapes already set on it.
Memory layout: reverse the dimensions
Core ML / MIL tensors are row-major; Julia arrays are column-major. A MIL tensor of shape [N, C, H, W] therefore has exactly the memory layout of a Julia Array of size (W, H, C, N). The graph functions here use that correspondence — they pass Julia's memory to BNNS as is, with no copy — so every argument is a Julia array whose size is the reverse of the model's shape (BNNSGraphArgument.size). In particular a MIL matrix [rows, cols] is a Julia (cols, rows) matrix, i.e. the transpose; use permutedims (or pass mil_order=true to bnns_graph_run) when you want model index order. BNNS graphs assume contiguous storage and ignore custom strides, so this is the only zero-copy mapping.
AppleAccelerate.bnns_graph_run — Function
bnns_graph_run(c, inputs...; func=nothing, mil_order=false, workspace=UInt8[]) -> Dict{String,Array}Run func on inputs (name => Array pairs, or a single Dict/NamedTuple), allocating the outputs, and return them keyed by output name.
g = AppleAccelerate.BNNSGraph("classifier.mlmodelc")
c = AppleAccelerate.BNNSGraphContext(g)
AppleAccelerate.bnns_graph_arguments(c) # names, eltypes, sizes
out = AppleAccelerate.bnns_graph_run(c, "image" => img)
probs = out["probabilities"]By default arrays use the zero-copy reversed-dimension layout (bnns_graph_arguments): pass size == BNNSGraphArgument.size. With mil_order=true inputs and outputs instead have the model's own shape (BNNSGraphArgument.shape) and index order — out[n, c, h, w] means what it means in the model — at the cost of one permutedims copy per array.
If the graph has dynamic dimensions, bind them first with bnns_graph_context_set_batch_size! or bnns_graph_context_set_dynamic_shapes!; an output whose shape is still unknown throws.
AppleAccelerate.bnns_graph_run! — Function
bnns_graph_run!(c, outputs, inputs; func=nothing, workspace=UInt8[]) -> outputsRun func of the graph behind context c, reading inputs and writing into the preallocated outputs. Both are collections of name => Array (a pair, a vector or tuple of pairs, a Dict, or a NamedTuple); together they must supply every argument of the function exactly once. Each array must be a dense Array whose element type and size match bnns_graph_arguments — note the reversed-dimension layout described there. Nothing is copied: BNNS reads and writes the arrays' memory directly.
Pass a reusable page-aligned workspace from bnns_graph_workspace to keep repeated calls allocation-free. A context must not be run from two threads at once.
AppleAccelerate.bnns_graph_workspace — Function
bnns_graph_workspace(c; func=nothing) -> Vector{UInt8}A page-aligned scratch buffer of the size func currently needs (bnns_graph_context_workspace_size), for the workspace keyword of bnns_graph_run! / bnns_graph_execute!. BNNSGraphContextExecute requires page alignment, which an ordinary Vector{UInt8} does not guarantee. Reusing one buffer across calls makes execution allocation-free on the BNNS side; without one BNNS allocates its own scratch on every call. The memory is released when the vector is garbage collected.
AppleAccelerate.bnns_graph_context_workspace_size — Function
bnns_graph_context_workspace_size(c, func=nothing) -> IntWorkspace size (bytes) required to execute func (BNNSGraphContextGetWorkspaceSize). Query it again after changing the batch size or dynamic shapes. Allocate a suitable buffer with bnns_graph_workspace.
AppleAccelerate.bnns_graph_context_set_batch_size! — Function
bnns_graph_context_set_batch_size!(c, n; func=nothing) -> cSet the batch size of a graph whose only dynamic dimension is a shared leading (MIL-order) batch dimension (BNNSGraphContextSetBatchSize). For anything more general use bnns_graph_context_set_dynamic_shapes!.
AppleAccelerate.bnns_graph_context_set_dynamic_shapes! — Function
bnns_graph_context_set_dynamic_shapes!(c, shapes; func=nothing) -> Vector{Pair{String,Dims}}Bind concrete input shapes to a graph compiled with dynamic dimensions (BNNSGraphContextSetDynamicShapes). shapes is a collection of name => dims pairs for (some of) the graph's inputs; dims is given the way Julia sees the array, i.e. size(A) of the array you will pass to bnns_graph_run (the reverse of the MIL shape — see bnns_graph_arguments). Inputs that are not mentioned keep the model's default shape.
Returns the resulting shape of every argument, in execute order, again in Julia order. A 0 in an output shape means BNNS cannot bound that dimension from the input shapes alone (it depends on input values).
AppleAccelerate.bnns_graph_argument_names — Function
Argument names of func (BNNSGraphGetArgumentNames).
AppleAccelerate.bnns_graph_execute! — Function
bnns_graph_execute!(c, arguments::Vector{bnns_graph_argument_t}; func=nothing, workspace=UInt8[]) -> cLow-level execute (BNNSGraphContextExecute): run func with raw argument buffers, ordered as bnns_graph_argument_names reports them (outputs first). The caller must keep the memory behind every argument alive for the call. Prefer bnns_graph_run / bnns_graph_run!, which build and validate the arguments from Julia arrays.
workspace must be page-aligned — get one from bnns_graph_workspace; leave it empty to let BNNS allocate its own scratch.
The compile-options accessors (bnns_compile_options_set_single_thread!, …_set_optimization!, …_set_output_path!, and their getters) and the remaining graph introspection helpers (bnns_graph_input_count, bnns_graph_input_names, bnns_graph_argument_intents, bnns_graph_argument_position, …) round out the family.
bnns_graph.h redirects BNNSGraphCompileFromFile, BNNSGraphContextExecute and eight other functions to _v2 symbols with an __asm__ label. The un-suffixed symbols that libBNNS still exports have a different, pre-release argument list, so the generated LibAccelerate.BNNSGraphCompileFromFile &c. must not be called directly (they crash); the wrappers on this page bind the _v2 symbols.
What's left to the raw layer
Everything not wrapped above is reachable through the raw AppleAccelerate.LibAccelerate layer. It falls into three groups:
- Deprecated classic tensor / DirectApply kernels (macOS 15 / iOS 18) —
BNNSMatMul, the activation-filter path,BNNSTile/BNNSTileBackward,BNNSCompareTensor,BNNSBandPart,BNNSGather/BNNSScatter(and their ND forms),BNNSShuffle, the clip family (BNNSClipByValue/…ByNorm/…ByGlobalNorm),BNNSComputeNorm,BNNSOptimizerStep, and theBNNSDirectApply{ActivationBatch,BroadcastMatMul,Quantizer}kernels. Superseded by the BNNS Graph API; intentionally not given an idiomatic wrapper. - The deprecated classic filter/layer API — the
BNNSFilterCreate*/BNNSFilterCreateLayer*constructors and their*FilterApply*/BNNSFusedFilterApply*execute paths (including the two-input / fused / loss / normalization / pooling / permute batch variants). - Exotic, training-only entry points that need training caches or opaque multi-kilobyte parameter blocks that cannot be validated generically: multi-head attention (
BNNSApplyMultiheadAttentionand its backward), the LSTM training-cache path (BNNSComputeLSTMTrainingCacheCapacity,BNNSDirectApplyLSTMBatchTrainingCaching/…Backward),BNNSComputeNormBackward, image crop/resize (BNNSCropResize/BNNSCropResizeBackward), and the fully-connected sparsification helpers (BNNSNDArrayFullyConnectedSparsifySparse{COO,CSR}).