Neural Network Primitives (BNNS)

AppleAccelerate wraps the current, non-deprecated slice of Apple's BNNS (Basic Neural Network Subroutines) library — 61 of the ~136 BNNS* C entry points. The bulk of the remainder are APIs Apple deprecated in macOS 15 / iOS 18: the classic filter/layer construction API and the deprecated classic + DirectApply tensor kernels (BNNSMatMul, BNNSTile, BNNSGather/BNNSScatter, the clip / norm family, BNNSOptimizerStep, …). Those are intentionally not wrapped — target the BNNS Graph API instead, which this page covers end to end: compile a Core ML model, inspect its arguments, and run inference on Julia arrays. Numerically verified helpers (transpose, copy, reductions, top-k, random generation, nearest neighbors, graph execution) are cross-checked against plain-Julia references; the remaining thin wrappers expose the rest of the current surface with exact FFI signatures for callers who need them.

Element types

Float32 works everywhere. Float16 and the integer / Bool types are accepted wherever the underlying kernel was verified at runtime to compute correctly — BNNS returns a success status with wrong values for some unsupported types, so each wrapper restricts its element types by dispatch:

FunctionElement types
bnns_transpose, same-type bnns_copy!Float16, Float32, Int8–Int64, UInt8–UInt64, Bool
converting bnns_copy!Float16 ↔ Float32, Int8/Int16/Int32/UInt8 → Float32, Float32 → Int32 (anything else throws)
bnns_reduceFloat16, Float32; Int32 for the integer-exact reductions
bnns_topkFloat16, Float32, Int8, Int16, Int32, UInt8, UInt16
bnns_in_topkFloat16, Float32
uniform / normal / categorical random fillsFloat16, Float32
bnns_random_fill_uniform_int!Int8–Int64, UInt8–UInt64
graph inputs / outputswhatever the model declares: Float16, Float32, integers, Bool
Namespace

These functions are not exported. Access them via the AppleAccelerate. prefix (e.g. AppleAccelerate.bnns_reduce).

Deprecated APIs are excluded

Apple deprecated the classic BNNS filter/layer API and much of the classic tensor/DirectApply surface (macOS 15 / iOS 18) in favour of the newer BNNS Graph API (BNNSGraph). This package does not wrap any of those deprecated entry points; use the Graph API for that functionality. The "What's left to the raw layer" section below lists the full excluded set.

Descriptors

BNNSArray builds a GC-safe BNNSNDArrayDescriptor view of a dense, contiguous Julia array. Internally the N-D op wrappers map a column-major Julia Array onto a BNNSDataLayout{N}DLastMajor descriptor with explicit strides, so BNNS axis k corresponds to Julia dimension k+1 (axis 0 is the contiguous/fastest axis).

AppleAccelerate.BNNSArray — Type
BNNSArray(A::AbstractArray)

A GC-safe BNNSNDArrayDescriptor view of a Julia array A, suitable for passing to BNNS routines via Ref. The wrapper keeps a reference to the backing array so it is not collected while the descriptor is alive; pass the underlying descriptor with Base.cconvert/Ref only inside a GC.@preserve block guarding A.

BNNS descriptors are layout aware. Julia stores arrays in column-major order, so this constructor reports the array using a column-major-friendly BNNS layout:

  • 1D Vector -> BNNSDataLayoutVector
  • 2D Matrix -> BNNSDataLayoutColumnMajorMatrix (BNNS size = (rows, cols)).

Only Float16, Float32 and Int32 dense, contiguous arrays are supported here; other element types or strided/transposed arrays should use the raw LibAccelerate layer directly.

source

Tensor manipulation

Stateless tensor ops that remain current, cross-validated against permutedims and plain copies.

FunctionMeaning
bnns_transposeswap two axes
bnns_copy!copy, optionally converting the element type
julia> M = Float32[1 2 3; 4 5 6];

julia> AppleAccelerate.bnns_transpose(M, 1, 2)
3×2 Matrix{Float32}:
 1.0  4.0
 2.0  5.0
 3.0  6.0

julia> AppleAccelerate.bnns_copy!(zeros(Float32, 2, 3), M) == M
true

julia> AppleAccelerate.bnns_copy!(zeros(Float16, 2, 3), M)     # converting copy: Float32 -> Float16
2×3 Matrix{Float16}:
 1.0  2.0  3.0
 4.0  5.0  6.0
AppleAccelerate.bnns_transpose — Function
bnns_transpose(A::Array, dim0, dim1) -> Array

Swap Julia dimensions dim0 and dim1 of A (1-based) via BNNSTranspose, equivalent to a permutedims that exchanges those two axes. Works for every element type BNNS can describe: Float16, Float32, Int8–Int64, UInt8–UInt64 and Bool.

source
AppleAccelerate.bnns_copy! — Function
bnns_copy!(dest::Array, src::Array) -> dest

Copy src into the equally-sized dest via BNNSCopy. For equal element types this is a plain element copy.

dest and src may have different element types, in which case BNNS converts: Float16 ↔ Float32 (the usual way to move data in and out of a half-precision graph), Int8/Int16/Int32/UInt8 → Float32, and Float32 → Int32. Any other pair throws an ArgumentError without calling BNNS.

source

Reductions

AppleAccelerate.bnns_reduce — Function
bnns_reduce(func::Symbol, input::Array; dim=1) -> Array

Reduce input along Julia dimension dim with func (:sum, :mean, :max, :min, :sumsquare, :l1, :l2, :product, :logsumexp) via BNNSDirectApplyReduction. The reduced axis collapses to length 1 and the result has the element type of input.

Supported element types: Float32, Float16 (computed in half precision, so sums saturate at floatmax(Float16) = 65504) and Int32. For Int32 only the reductions that are exact in integers are offered (:sum, :max, :min, :sumsquare, :l1, :product); :mean, :l2 and :logsumexp throw.

source

DirectApply kernels

Fused kernels that run without an explicit filter handle.

AppleAccelerate.bnns_topk — Function
bnns_topk(input::Array, K; dim=1) -> (values, indices)

Top-K values and their 0-based indices along Julia dimension dim via BNNSDirectApplyTopK. values has the element type of input, indices is Int32. Comparable to sort-based partialsortperm per slice.

Supported element types: Float32, Float16, Int8, Int16, Int32, UInt8, UInt16 (BNNS rejects the wider integer types).

source
AppleAccelerate.bnns_in_topk — Function
bnns_in_topk(input::Array, targets::Array{Int32}, K; dim=1) -> Array{Bool}

For each batch column, test whether the targets class index is among the top-K scores of input along Julia dimension dim (BNNSDirectApplyInTopK). input may be Float32 or Float16.

source

Utility queries

AppleAccelerate.bnns_layout_rank — Function
bnns_layout_rank(layout::BNNSDataLayout) -> Int

Rank (number of dimensions) encoded by a BNNSDataLayout constant, via BNNSDataLayoutGetRank.

source
AppleAccelerate.bnns_tensor_allocation_size — Function
bnns_tensor_allocation_size(A::Array) -> Int

Bytes required to allocate a BNNSTensor describing A (BNNSTensorGetAllocationSize). Uses the modern BNNSTensor struct (rank + shape/stride), distinct from the legacy BNNSNDArrayDescriptor.

source

Random number generation

BNNSRandomGenerator is an AES-CTR generator with an optional seed; the fill functions populate arrays in place and the state can be snapshot and restored for reproducibility.

AppleAccelerate.BNNSRandomGenerator — Type
BNNSRandomGenerator([seed]) -> BNNSRandomGenerator

A BNNS random number generator handle (AES-CTR method). Construct with an optional 64-bit seed for reproducibility (BNNSCreateRandomGeneratorWithSeed, or BNNSCreateRandomGenerator when omitted). The handle is destroyed automatically by a finalizer (BNNSDestroyRandomGenerator).

Use with bnns_random_fill_uniform!, bnns_random_fill_normal!, bnns_random_fill_uniform_int!, bnns_random_fill_categorical! and the bnns_random_state/bnns_random_state! round-trip.

source
AppleAccelerate.bnns_random_fill_uniform! — Function
bnns_random_fill_uniform!(g::BNNSRandomGenerator, A::Array, lo=0f0, hi=1f0) -> A

Fill A (Float32 or Float16) with i.i.d. uniform samples on [lo, hi) (BNNSRandomFillUniformFloat). For Float16 the samples are rounded to half precision, so a value can round up to exactly hi.

source
AppleAccelerate.bnns_random_fill_uniform_int! — Function
bnns_random_fill_uniform_int!(g::BNNSRandomGenerator, A::Array{<:Integer}, lo, hi) -> A

Fill integer array A with i.i.d. uniform samples on the half-open range [lo, hi) (BNNSRandomFillUniformInt). A may be Int8, Int16, Int32, Int64, UInt8, UInt16, UInt32 or UInt64; the range must fit the element type.

source
AppleAccelerate.bnns_random_fill_categorical! — Function
bnns_random_fill_categorical!(g::BNNSRandomGenerator, out::Array{T}, probs::Array{T}; log_probs=false) -> out

Draw categorical samples (0-based category indices, stored as floating point) into out using per-category weights probs (BNNSRandomFillCategoricalFloat). Pass log_probs=true if probs holds log probabilities. T is Float32 or Float16; out and probs must share it (BNNS silently mis-samples mixed precisions).

source

Nearest neighbors

AppleAccelerate.BNNSNearestNeighbors — Type
BNNSNearestNeighbors(max_samples, n_features, n_neighbors; T=Float32) -> BNNSNearestNeighbors

A brute-force k-nearest-neighbours index (BNNSCreateNearestNeighbors) holding up to max_samples reference points of dimension n_features, answering n_neighbors-NN queries. Destroyed automatically (BNNSDestroyNearestNeighbors).

Add reference points with bnns_knn_load! and query with bnns_knn_query.

source
AppleAccelerate.bnns_knn_load! — Function
bnns_knn_load!(knn::BNNSNearestNeighbors, data::Matrix{Float32}) -> Int

Append reference samples to the index (BNNSNearestNeighborsLoad). data is n_features × n_new_samples (each column is one sample, matching BNNS's feature-major layout). Returns the number of samples loaded.

source
AppleAccelerate.bnns_knn_query — Function
bnns_knn_query(knn::BNNSNearestNeighbors, sample_number) -> (indices, distances)

Return the n_neighbors nearest reference points to the (0-based) loaded sample sample_number (BNNSNearestNeighborsGetInfo): their 0-based indices (Vector{Int32}) and Float32 distances.

source

BNNS Graph API

The modern, non-deprecated pipeline (macOS 15+): compile a Core ML model into a BNNSGraph, make an executable BNNSGraphContext, look at what it expects with bnns_graph_arguments, and run it on Julia arrays with bnns_graph_run / bnns_graph_run!.

The input is a compiled Core ML model — the .mlmodelc directory that Xcode or xcrun coremlcompiler compile model.mlpackage out/ produces (ML Program models only). There is no in-memory graph builder in this API. A .mlmodelc is a directory holding a textual MIL program, model.mil, plus a weights blob; the example below writes a tiny one by hand so that it is self-contained.

modeldir = joinpath(mktempdir(), "dense.mlmodelc"); mkpath(modeldir)
write(joinpath(modeldir, "model.mil"), """
program(1.3)
[buildInfo = dict<string, string>({{"coremlc-component-MIL", "handwritten"}})]
{
    func main<ios16>(tensor<fp32, [2, 3]> x) {
            tensor<fp32, [3]> b = const()[name = string("b"), val = tensor<fp32, [3]>([1.0, -2.0, 0.5])];
            tensor<fp32, [2, 3]> s = add(x = x, y = b)[name = string("s")];
            tensor<fp32, [2, 3]> z = relu(x = s)[name = string("z")];
        } -> (z);
}
""")

graph = AppleAccelerate.BNNSGraph(modeldir)
ctx   = AppleAccelerate.BNNSGraphContext(graph)
AppleAccelerate.bnns_graph_arguments(ctx)
2-element Vector{AppleAccelerate.BNNSGraphArgument}:
 BNNSGraphArgument("z", :out, Float32, shape=(2, 3), size=(3, 2))
 BNNSGraphArgument("x", :in, Float32, shape=(2, 3), size=(3, 2))

Memory layout

MIL tensors are row-major, Julia arrays are column-major, so a model tensor of shape [2, 3] has the memory layout of a Julia array of size (3, 2). The graph functions pass Julia's memory to BNNS untouched, which means every argument is a Julia array whose size is the reverse of the model's shape — that is the size field reported above. For an image model, [N, C, H, W] is a Julia (W, H, C, N) array. BNNS graphs ignore custom strides, so this is the only zero-copy mapping.

X = Float32[1 2 3; -4 0 6]                       # in the model's [2, 3] index order
out = AppleAccelerate.bnns_graph_run(ctx, "x" => permutedims(X))   # (3, 2): reversed dims
@assert permutedims(out["z"]) == max.(X .+ Float32[1 -2 0.5], 0)

# mil_order=true does that permutedims for you, on the way in and on the way out
out = AppleAccelerate.bnns_graph_run(ctx, "x" => X; mil_order = true)
@assert out["z"] == max.(X .+ Float32[1 -2 0.5], 0)

For repeated inference preallocate the outputs and a page-aligned workspace, and nothing is allocated per call on the BNNS side:

Z  = zeros(Float32, 3, 2)
ws = AppleAccelerate.bnns_graph_workspace(ctx)
AppleAccelerate.bnns_graph_run!(ctx, "z" => Z, "x" => permutedims(X); workspace = ws)
@assert permutedims(Z) == out["z"]

A half-precision model takes and returns Float16 arrays; convert with bnns_copy! (or plain Float16.(x)). Models with a dynamic batch dimension are bound with bnns_graph_context_set_batch_size!, more general dynamic shapes with bnns_graph_context_set_dynamic_shapes!. A context carries mutable state and must be used by one thread at a time; make one context per task.

AppleAccelerate.BNNSGraphCompileOptions — Type
BNNSGraphCompileOptions(; single_thread=nothing, generate_debug_info=nothing,
                          optimization=nothing, log_mask=nothing,
                          output_path=nothing, output_fd=nothing) -> BNNSGraphCompileOptions

Options controlling BNNSGraphCompileFromFile, backed by BNNSGraphCompileOptionsMakeDefault and destroyed by a finalizer (BNNSGraphCompileOptionsDestroy). Any keyword left nothing keeps the BNNS default. optimization is :performance or :ir_size. Individual fields can also be read/written with the accessor functions below.

source
AppleAccelerate.BNNSGraph — Type
BNNSGraph(filename; func=nothing, options=BNNSGraphCompileOptions()) -> BNNSGraph

Compile the compiled Core ML model (.mlmodelc directory, or the model.mil inside it) at filename — optionally only the named func inside it — into an executable graph via BNNSGraphCompileFromFile. Throws if BNNS cannot compile the model (unsupported op, malformed program, missing file). The returned handle feeds BNNSGraphContext and the graph-introspection helpers, and the compiled graph's memory is released by a finalizer once the graph and every context made from it are unreachable.

.mlmodelc is what Xcode / xcrun coremlcompiler compile produce from an .mlpackage; only ML Program models (MIL), not the older NeuralNetwork format, are accepted by BNNS.

source
AppleAccelerate.BNNSGraphContext — Type
BNNSGraphContext(g::BNNSGraph) -> BNNSGraphContext

An executable context for a compiled BNNSGraph (BNNSGraphContextMake), destroyed by a finalizer (BNNSGraphContextDestroy). It holds the mutable execution state (dynamic shapes, streaming state), keeps its graph alive, and must be used by one thread at a time; make one context per task for concurrent inference. Feed it to bnns_graph_run / bnns_graph_run!, or to the low-level bnns_graph_execute!.

source
AppleAccelerate.BNNSGraphArgument — Type
BNNSGraphArgument

Description of one argument of a graph function, as returned by bnns_graph_arguments:

  • name::String
  • intent::Symbol — :in, :out or :inout
  • eltype::DataType — Julia element type (Float16, Float32, Int32, Bool, …)
  • shape::Dims — the shape as written in the model (MIL / row-major order)
  • size::Dims — size of the Julia Array to pass for it: reverse(shape)

A dimension of 0 (in either tuple) is dynamic and not bound yet.

source
AppleAccelerate.bnns_graph_arguments — Function
bnns_graph_arguments(c::BNNSGraphContext; func=nothing) -> Vector{BNNSGraphArgument}
bnns_graph_arguments(g::BNNSGraph; func=nothing)

Names, intents, element types and shapes of every argument of func, in execute order (outputs first). Given a context, shapes reflect any batch size / dynamic shapes already set on it.

Memory layout: reverse the dimensions

Core ML / MIL tensors are row-major; Julia arrays are column-major. A MIL tensor of shape [N, C, H, W] therefore has exactly the memory layout of a Julia Array of size (W, H, C, N). The graph functions here use that correspondence — they pass Julia's memory to BNNS as is, with no copy — so every argument is a Julia array whose size is the reverse of the model's shape (BNNSGraphArgument.size). In particular a MIL matrix [rows, cols] is a Julia (cols, rows) matrix, i.e. the transpose; use permutedims (or pass mil_order=true to bnns_graph_run) when you want model index order. BNNS graphs assume contiguous storage and ignore custom strides, so this is the only zero-copy mapping.

source
AppleAccelerate.bnns_graph_run — Function
bnns_graph_run(c, inputs...; func=nothing, mil_order=false, workspace=UInt8[]) -> Dict{String,Array}

Run func on inputs (name => Array pairs, or a single Dict/NamedTuple), allocating the outputs, and return them keyed by output name.

g = AppleAccelerate.BNNSGraph("classifier.mlmodelc")
c = AppleAccelerate.BNNSGraphContext(g)
AppleAccelerate.bnns_graph_arguments(c)          # names, eltypes, sizes
out = AppleAccelerate.bnns_graph_run(c, "image" => img)
probs = out["probabilities"]

By default arrays use the zero-copy reversed-dimension layout (bnns_graph_arguments): pass size == BNNSGraphArgument.size. With mil_order=true inputs and outputs instead have the model's own shape (BNNSGraphArgument.shape) and index order — out[n, c, h, w] means what it means in the model — at the cost of one permutedims copy per array.

If the graph has dynamic dimensions, bind them first with bnns_graph_context_set_batch_size! or bnns_graph_context_set_dynamic_shapes!; an output whose shape is still unknown throws.

source
AppleAccelerate.bnns_graph_run! — Function
bnns_graph_run!(c, outputs, inputs; func=nothing, workspace=UInt8[]) -> outputs

Run func of the graph behind context c, reading inputs and writing into the preallocated outputs. Both are collections of name => Array (a pair, a vector or tuple of pairs, a Dict, or a NamedTuple); together they must supply every argument of the function exactly once. Each array must be a dense Array whose element type and size match bnns_graph_arguments — note the reversed-dimension layout described there. Nothing is copied: BNNS reads and writes the arrays' memory directly.

Pass a reusable page-aligned workspace from bnns_graph_workspace to keep repeated calls allocation-free. A context must not be run from two threads at once.

source
AppleAccelerate.bnns_graph_workspace — Function
bnns_graph_workspace(c; func=nothing) -> Vector{UInt8}

A page-aligned scratch buffer of the size func currently needs (bnns_graph_context_workspace_size), for the workspace keyword of bnns_graph_run! / bnns_graph_execute!. BNNSGraphContextExecute requires page alignment, which an ordinary Vector{UInt8} does not guarantee. Reusing one buffer across calls makes execution allocation-free on the BNNS side; without one BNNS allocates its own scratch on every call. The memory is released when the vector is garbage collected.

source
AppleAccelerate.bnns_graph_context_set_dynamic_shapes! — Function
bnns_graph_context_set_dynamic_shapes!(c, shapes; func=nothing) -> Vector{Pair{String,Dims}}

Bind concrete input shapes to a graph compiled with dynamic dimensions (BNNSGraphContextSetDynamicShapes). shapes is a collection of name => dims pairs for (some of) the graph's inputs; dims is given the way Julia sees the array, i.e. size(A) of the array you will pass to bnns_graph_run (the reverse of the MIL shape — see bnns_graph_arguments). Inputs that are not mentioned keep the model's default shape.

Returns the resulting shape of every argument, in execute order, again in Julia order. A 0 in an output shape means BNNS cannot bound that dimension from the input shapes alone (it depends on input values).

source
AppleAccelerate.bnns_graph_execute! — Function
bnns_graph_execute!(c, arguments::Vector{bnns_graph_argument_t}; func=nothing, workspace=UInt8[]) -> c

Low-level execute (BNNSGraphContextExecute): run func with raw argument buffers, ordered as bnns_graph_argument_names reports them (outputs first). The caller must keep the memory behind every argument alive for the call. Prefer bnns_graph_run / bnns_graph_run!, which build and validate the arguments from Julia arrays.

workspace must be page-aligned — get one from bnns_graph_workspace; leave it empty to let BNNS allocate its own scratch.

source

The compile-options accessors (bnns_compile_options_set_single_thread!, …_set_optimization!, …_set_output_path!, and their getters) and the remaining graph introspection helpers (bnns_graph_input_count, bnns_graph_input_names, bnns_graph_argument_intents, bnns_graph_argument_position, …) round out the family.

Versioned symbols

bnns_graph.h redirects BNNSGraphCompileFromFile, BNNSGraphContextExecute and eight other functions to _v2 symbols with an __asm__ label. The un-suffixed symbols that libBNNS still exports have a different, pre-release argument list, so the generated LibAccelerate.BNNSGraphCompileFromFile &c. must not be called directly (they crash); the wrappers on this page bind the _v2 symbols.

What's left to the raw layer

Everything not wrapped above is reachable through the raw AppleAccelerate.LibAccelerate layer. It falls into three groups:

  • Deprecated classic tensor / DirectApply kernels (macOS 15 / iOS 18) — BNNSMatMul, the activation-filter path, BNNSTile/BNNSTileBackward, BNNSCompareTensor, BNNSBandPart, BNNSGather/BNNSScatter (and their ND forms), BNNSShuffle, the clip family (BNNSClipByValue/…ByNorm/ …ByGlobalNorm), BNNSComputeNorm, BNNSOptimizerStep, and the BNNSDirectApply{ActivationBatch,BroadcastMatMul,Quantizer} kernels. Superseded by the BNNS Graph API; intentionally not given an idiomatic wrapper.
  • The deprecated classic filter/layer API — the BNNSFilterCreate* / BNNSFilterCreateLayer* constructors and their *FilterApply* / BNNSFusedFilterApply* execute paths (including the two-input / fused / loss / normalization / pooling / permute batch variants).
  • Exotic, training-only entry points that need training caches or opaque multi-kilobyte parameter blocks that cannot be validated generically: multi-head attention (BNNSApplyMultiheadAttention and its backward), the LSTM training-cache path (BNNSComputeLSTMTrainingCacheCapacity, BNNSDirectApplyLSTMBatchTrainingCaching / …Backward), BNNSComputeNormBackward, image crop/resize (BNNSCropResize / BNNSCropResizeBackward), and the fully-connected sparsification helpers (BNNSNDArrayFullyConnectedSparsifySparse{COO,CSR}).