Architecture
This page explains how AppleAccelerate is structured internally, and why it integrates with the rest of the ecosystem through its own namespace rather than by hooking other packages' interfaces.
Two-layer design
AppleAccelerate is built in two layers:
Raw ABI layer —
AppleAccelerate.LibAccelerate. This submodule (src/lib/LibAccelerate.jl) is auto-generated with Clang.jl directly from Apple's Accelerate C headers. It is a near 1:1 mirror of the C API — roughly 1400@ccallwrappers plus the matching structs and enums — and is committed but never hand-edited (regenerate it withjulia --project=gen gen/generate.jl). Because Accelerate's struct- and enum-heavy subframeworks (Quadrature, Sparse, BNNS, vImage) are painful and error-prone to bind by hand, generating this layer keeps every field offset and enum value in sync with the SDK automatically.Idiomatic layer — the hand-written
src/*.jlfiles. These provide the ergonomic Julia API (AppleAccelerate.exp,integrate,bnns_reduce,AASparseMatrix, …). They own the niceties — Julia functions, keyword options, tidy return values, broadcasting — and call intoLibAcceleratefor the actual ABI work, never naming a C field offset or enum integer themselves.
Most users only ever touch the idiomatic layer. When you need a symbol that has no idiomatic wrapper yet, the raw layer is available directly:
using AppleAccelerate
AppleAccelerate.LibAccelerate.some_unwrapped_symbol(args...)No package extensions
AppleAccelerate ships no package extensions, and defines no methods on other packages' functions. Everything it provides lives under the AppleAccelerate. namespace, and the package exports nothing.
This is a deliberate choice. Earlier versions did hook AbstractFFTs.jl and NNlib.jl through extensions, which meant that merely loading AppleAccelerate — even for an unrelated feature like BLAS forwarding — silently changed what fft, plan_fft, or batched_mul! did elsewhere in the session. Because Accelerate's kernels only cover a subset of the input space (vDSP's FFT is power-of-2 and 1-D/2-D only), those methods won dispatch and then failed on inputs the general backend would have handled fine. See issue #139.
The consequence for you is simple and predictable:
using AppleAccelerate, FFTW
fft(x) # always FFTW — loading AppleAccelerate changes nothing
AppleAccelerate.fft(x) # explicitly vDSP, for power-of-2 1-D/2-D inputYou choose Accelerate per call site, by name. Nothing is intercepted, and benchmarking one against the other is a matter of changing the prefix.
Why LinearAlgebra and SparseArrays are hard dependencies
LinearAlgebra and SparseArrays are regular [deps], not weak dependencies:
- They are standard libraries that are effectively always present in the Julia sysimage, so making them weak would save nothing.
- BLAS/LAPACK forwarding needs
LinearAlgebraat load time (__init__installs Accelerate into libblastrampoline), so it cannot be deferred. - Keeping
SparseArraysa hard dependency letsAAFactorizationsubtypeLinearAlgebra.Factorizationand letsAASparseMatrixinteroperate withSparseMatrixCSCdirectly.
BLAS/LAPACK forwarding is the one place AppleAccelerate does change global behavior — that is the documented purpose of loading it, and it goes through libblastrampoline, a mechanism designed for exactly this.
Coverage of the raw layer
How much of the generated layer has an idiomatic wrapper is audited exactly by gen/coverage_audit.jl (it inspects lowered IR rather than grepping for C names). The vDSP result, including the handful of functions deliberately left to the raw layer:
AppleAccelerate.VDSP_COVERAGE — Constant
AppleAccelerate.VDSP_COVERAGEDocumentation of idiomatic-wrapper coverage for vDSP.h. Companion to AppleAccelerate.VMATH_COVERAGE.
The numbers come from gen/coverage_audit.jl, which inspects the lowered IR of every method in the package and lists the generated LibAccelerate.vDSP_* functions that nothing calls. Do not estimate coverage by grepping src/ for C symbol names: most wrappers assemble the name at macro-expansion time (Symbol(string("vDSP_vfix", intname, suff))), so whole families that are fully wrapped — the 48 vfix*/vflt* conversions, the fixed-point _s1_15/_s8_24 kernels, every D-suffixed Float64 variant — look "missing" to a text search.
Tally (macOS 26 SDK raw layer: 469 vDSP_* functions)
| Status | Count | Functions |
|---|---|---|
| Reached by an idiomatic wrapper | 465 | everything not listed below |
| Intentionally left to the raw layer | 4 | vDSP_DFT_CreateSetup, vDSP_DFT_zop, vDSP_fft2d_zrip, vDSP_fft2d_zripD |
| Remaining | 0 | — |
The raw layer itself captures ~93% of vDSP.h; the functions libclang drops are the overloads whose signatures use arm_neon/simd register types (see gen/README.md), which have no array surface.
Intentionally left to the raw layer
vDSP_DFT_CreateSetup/vDSP_DFT_zop— the original DFT entry points. The header says "We recommend you usevDSP_DFT_zop_CreateSetupinstead of this routine"; that replacement family (vDSP_DFT_zop_CreateSetup(D)+vDSP_DFT_Execute(D)) is whatplan_dft/dftwrap. The old pair is Float32-only and computes the identical transform.vDSP_fft2d_zrip(D)— the in-place, no-temporary 2-D real FFT. The same transform is wrapped throughvDSP_fft2d_zrop(D)(the matrix methods ofrfft) and the caller-workspace variantsvDSP_fft2d_zropt(D)/vDSP_fft2d_zript(D). A JuliaMatrix{T}must be repacked into vDSP's even/odd split-complex layout before any of these can run, so "in place on the packed buffer" has no Julian surface of its own, and thezriptform (caller-owned scratch instead of an internalmalloc) is the one worth exposing.
Closed by this file
vDSP_biquadm_SetCoefficientsDouble,vDSP_biquadm_SetCoefficientsSingleD,vDSP_biquadm_SetTargetsDouble,vDSP_biquadm_SetTargetsSingleD— cross-precision methods ofbiquadm_setcoefficients!/biquadm_settargets!.vDSP_FFT16_copv,vDSP_FFT32_copv— allocation-freefft16!,bfft16!,fft32!,bfft32!.