FFT & Transforms (vDSP)
AppleAccelerate wraps Apple's vDSP transform functions for FFT, real FFT, DFT, and DCT.
FFT
The FFT API wraps Apple's vDSP FFT functions and follows the same naming conventions as AbstractFFTs.jl. Both 1D vectors and 2D matrices are supported with ComplexF64 and ComplexF32 inputs. A 1D fft/ifft/bfft/rfft/irfft (called without an explicit plan_fft setup) accepts any power-of-two length, plus non-power-of-two lengths of the form f * 2^k with f ∈ {3, 5, 15} — for complex transforms k ≥ 3 (so the smallest such lengths are 24, 40, 120), and for real transforms k ≥ 4. Powers of two use vDSP's fast FFT path, while the mixed-radix lengths transparently use Apple's DFT (see is_supported_fft_length). 2D transforms remain power-of-2 only. Unsupported lengths throw an ArgumentError.
AppleAccelerate's FFT is a direct vDSP binding: 1-D transforms support powers of two plus f * 2^k lengths (f ∈ {3, 5, 15}, with k ≥ 3 for complex and k ≥ 4 for real transforms), while 2-D transforms are power-of-2 only. Only 1-D vectors and 2-D matrices are supported, and it is reached exclusively through the AppleAccelerate. prefix. It does not plug into the AbstractFFTs.jl interface — loading AppleAccelerate never changes what fft/plan_fft do in your session, and will never override FFTW.jl.
For general sizes, 3-D and higher arrays, dimension-selective (region) transforms, or the standard Julia FFT interface, use FFTW.jl. vDSP is mainly competitive at small transform sizes; FFTW is typically faster at larger ones.
Real FFT (rfft, irfft, brfft) is likewise available only through the AppleAccelerate. prefix.
x = randn(ComplexF64, 1024)
# Forward FFT (auto-creates plan)
X = AppleAccelerate.fft(x)
# Normalized inverse FFT
x_recovered = AppleAccelerate.ifft(X)
# Unnormalized inverse (backward) FFT
X_back = AppleAccelerate.bfft(X)
# Reusable plan for repeated transforms of the same size
setup = AppleAccelerate.plan_fft(x)
X = AppleAccelerate.fft(x, setup)
x_recovered = AppleAccelerate.ifft(X, setup)2D FFT works the same way — fft dispatches on the input shape:
x2 = randn(ComplexF64, 16, 32)
X2 = AppleAccelerate.fft(x2)
x2_recovered = AppleAccelerate.ifft(X2)Batched 1D FFT
Passing a dims argument (1 = each column, 2 = each row) computes a batch of independent 1D transforms over a matrix in a single call, like FFTW's fft(A, dims). This wraps vDSP_fftm_zop. The transform length (size(A, dims)) must be a power of 2:
A = randn(ComplexF64, 16, 8)
F = AppleAccelerate.fft(A, 1) # FFT of each column
A1 = AppleAccelerate.ifft(F, 1) # normalized inverse
@assert A1 ≈ AIn-place FFT
In-place variants avoid allocating the output array:
x = randn(ComplexF64, 1024)
y = copy(x)
AppleAccelerate.fft!(y) # forward, modifies y in-place
AppleAccelerate.ifft!(y) # inverse, modifies y in-place
@assert y ≈ xTemp-buffer variants
For repeated transforms of the same size, an FFTWorkspace supplies vDSP with a reusable scratch buffer (the *t vDSP variants), avoiding its internal per-call allocation. Pass it as the last argument to fft/fft!/bfft/ifft (1D, 2D, or batched) and to the real transforms. Results are identical to the non-workspace calls.
x = randn(ComplexF64, 1024)
setup = AppleAccelerate.plan_fft(x)
ws = AppleAccelerate.FFTWorkspace(x)
X = AppleAccelerate.fft(x, setup, ws)
@assert AppleAccelerate.ifft(X, setup, ws) ≈ xAppleAccelerate.plan_fft — Function
plan_fft(x::VecOrMat{Complex{T}}) where T <: Union{Float32, Float64}
plan_fft(n::Integer, [T=Float64], [radix=2])Create a reusable FFT setup object for repeated transforms of the same size. Wraps vDSP_create_fftsetup / vDSP_create_fftsetupD. The setup is automatically destroyed when garbage collected.
AppleAccelerate.fft — Function
fft(x::VecOrMat{Complex{T}}, [setup::FFTSetup{T}])Compute the forward FFT of x via Apple vDSP. Supports 1D vectors and 2D matrices with ComplexF32 or ComplexF64 elements.
For a 1D vector with no explicit setup, non-power-of-2 lengths of the form f * 2^k with f ∈ {3, 5, 15} and k ≥ 3 are also supported and transparently use Apple's mixed-radix DFT (see is_supported_fft_length); an unsupported length throws an ArgumentError. When a setup::FFTSetup is supplied, or for 2D inputs, all dimensions must be powers of 2. If setup is omitted, a temporary plan is created automatically.
The 1D no-setup method takes a fallback keyword: fft(x; fallback = FFTW.fft) returns fallback(x) instead of throwing when vDSP does not support length(x); it is never called for a supported length. bfft, ifft and rfft accept it likewise. For repeated transforms see fftplan. Wraps vDSP_fft_zop (1D) / vDSP_fft2d_zop (2D).
fft(x::Matrix{Complex{T}}, dims::Integer, [setup::FFTSetup{T}])Batched 1D forward FFT: transform each column (dims = 1) or each row (dims = 2) of x independently, like FFTW's fft(x, dims). The transform length (size(x, dims)) must be a power of 2. Wraps vDSP_fftm_zop.
fft(x, setup::FFTSetup, ws::FFTWorkspace)
fft!(x, setup::FFTSetup, ws::FFTWorkspace)Temp-buffer forward FFT: same as fft/fft! but uses the caller-supplied FFTWorkspace ws as scratch, avoiding vDSP's internal buffer allocation on each call. Applies to 1D vectors and 2D matrices; all dimensions must be powers of 2. Wraps vDSP_fft_zopt / vDSP_fft_zipt (1D) and vDSP_fft2d_zopt / vDSP_fft2d_zipt (2D).
fft(x::Matrix, dims, setup, ws) ; fft!(x::Matrix, dims, setup[, ws])Batched 1D FFT along dims (1 = columns, 2 = rows) with a preallocated FFTWorkspace ws (temp-buffer form) and/or in-place (fft!) operation. See fft. Wraps vDSP_fftm_zip / vDSP_fftm_zipt / vDSP_fftm_zopt.
AppleAccelerate.ifft — Function
ifft(x::VecOrMat{Complex{T}}, [setup::FFTSetup{T}])Compute the normalized inverse FFT of x via Apple vDSP. Equivalent to bfft(x) / length(x). Satisfies ifft(fft(x)) ≈ x.
For a 1D vector with no explicit setup, non-power-of-2 lengths of the form f * 2^k with f ∈ {3, 5, 15} and k ≥ 3 are also supported (via Apple's mixed-radix DFT); an unsupported length throws an ArgumentError. With an explicit setup::FFTSetup, or for 2D inputs, all dimensions must be powers of 2. Wraps vDSP_fft_zop (1D) / vDSP_fft2d_zop (2D).
ifft(x::Matrix{Complex{T}}, dims::Integer, [setup::FFTSetup{T}])Batched 1D normalized inverse FFT along dimension dims (1 = columns, 2 = rows). Satisfies ifft(fft(x, dims), dims) ≈ x. Wraps vDSP_fftm_zop.
ifft(x, setup::FFTSetup, ws::FFTWorkspace)
ifft!(x, setup::FFTSetup, ws::FFTWorkspace)Temp-buffer normalized inverse FFT (see ifft/ifft!), using ws as scratch.
ifft(x::Matrix, dims, setup, ws) ; ifft!(x::Matrix, dims, setup[, ws])Batched normalized inverse FFT along dims, temp-buffer and/or in-place. See ifft.
AppleAccelerate.bfft — Function
bfft(x::VecOrMat{Complex{T}}, [setup::FFTSetup{T}])Compute the unnormalized inverse (backward) FFT of x via Apple vDSP. The result is not divided by length(x); use ifft for the normalized version.
For a 1D vector with no explicit setup, non-power-of-2 lengths of the form f * 2^k with f ∈ {3, 5, 15} and k ≥ 3 are also supported (via Apple's mixed-radix DFT); an unsupported length throws an ArgumentError. With an explicit setup::FFTSetup, or for 2D inputs, all dimensions must be powers of 2. Wraps vDSP_fft_zop (1D) / vDSP_fft2d_zop (2D).
bfft(x::Matrix{Complex{T}}, dims::Integer, [setup::FFTSetup{T}])Batched 1D unnormalized inverse (backward) FFT along dimension dims (1 = columns, 2 = rows). Use ifft for the normalized version. Wraps vDSP_fftm_zop.
bfft(x, setup::FFTSetup, ws::FFTWorkspace)
bfft!(x, setup::FFTSetup, ws::FFTWorkspace)Temp-buffer unnormalized inverse FFT (see bfft/bfft!), using ws as scratch. Wraps the zopt/zipt vDSP variants.
bfft(x::Matrix, dims, setup, ws) ; bfft!(x::Matrix, dims, setup[, ws])Batched unnormalized inverse FFT along dims, temp-buffer and/or in-place. See bfft.
AppleAccelerate.fft! — Function
fft!(x::VecOrMat{Complex{T}}, [setup::FFTSetup{T}])Compute the forward FFT of x in-place via Apple vDSP, overwriting x with the result. Wraps vDSP_fft_zip (1D) / vDSP_fft2d_zip (2D).
AppleAccelerate.ifft! — Function
ifft!(x::VecOrMat{Complex{T}}, [setup::FFTSetup{T}])Compute the normalized inverse FFT of x in-place via Apple vDSP. Satisfies ifft!(fft!(copy(x))) ≈ x. Wraps vDSP_fft_zip (1D) / vDSP_fft2d_zip (2D).
AppleAccelerate.bfft! — Function
bfft!(x::VecOrMat{Complex{T}}, [setup::FFTSetup{T}])Compute the unnormalized inverse FFT of x in-place via Apple vDSP. Wraps vDSP_fft_zip (1D) / vDSP_fft2d_zip (2D).
AppleAccelerate.FFTWorkspace — Type
FFTWorkspace{T}(n::Integer)
FFTWorkspace(x::AbstractArray)Preallocated split-complex scratch buffer for the temp-buffer FFT variants (fft(x, setup, ws), fft!(x, setup, ws), rfft(x, setup, ws), …). Holds two Vector{T} (real/imag parts) of length n. Construct from an array to size it automatically (length(x) elements, which is always sufficient). Reusable across transforms of the same or smaller size.
AppleAccelerate.is_supported_fft_length — Function
is_supported_fft_length(n::Integer) -> BoolReturn true if a 1D complex FFT of length n is supported by Apple vDSP via the idiomatic fft/ifft/bfft API. Supported lengths are any power of two, plus f * 2^k with f ∈ {3, 5, 15} and k ≥ 3 (the smallest such non-power-of-two lengths are 24, 40, and 120).
Real FFT
rfft computes the FFT of real input, returning only the non-redundant complex coefficients (length N÷2+1). irfft inverts it back to real output.
x = randn(Float64, 1024)
X = AppleAccelerate.rfft(x) # Complex vector of length 513
x_recovered = AppleAccelerate.irfft(X, 1024) # Back to real, length 1024
@assert x_recovered ≈ x1D real transforms called without an explicit setup also accept non-power-of-2 lengths of the form f * 2^k with f ∈ {3, 5, 15} and k ≥ 4 (e.g. 48, 160, 240), via Apple's real-input mixed-radix DFT (vDSP_DFT_zrop_CreateSetup).
2D real FFT is supported for power-of-2 dimensions and matches FFTW's rfft layout — an n1×n2 real matrix transforms to an (n1÷2+1)×n2 complex matrix:
x2 = randn(Float64, 16, 32)
X2 = AppleAccelerate.rfft(x2) # 9×32 complex matrix
x2_recovered = AppleAccelerate.irfft(X2, 16) # back to real, 16×32
@assert x2_recovered ≈ x2Batched real FFT (rfft(x, dims) / irfft(X, n, dims)) transforms each column (dims = 1) or row (dims = 2) of a real matrix independently, like FFTW's rfft(x, dims). In-place real variants (rfft!) reuse the input buffer as vDSP's scratch (overwriting it), and workspace variants accept an FFTWorkspace.
AppleAccelerate.plan_rfft — Function
plan_rfft(x::VecOrMat{T}) where T <: Union{Float32, Float64}Create a reusable FFT setup for real-input forward transforms of the same size as x. For a matrix, the setup covers both dimensions (and can also be reused for the matching inverse transforms). Wraps vDSP_create_fftsetup.
AppleAccelerate.rfft — Function
rfft(x::VecOrMat{T}, [setup::FFTSetup{T}])Compute the forward FFT of a real-valued vector or matrix x via Apple vDSP. Returns the non-redundant complex coefficients: length N÷2+1 for a vector of length N, and size (n1÷2+1)×n2 for an n1×n2 matrix (matching FFTW's rfft).
For a 1D vector with no explicit setup, non-power-of-2 lengths supported by Apple's real-input mixed-radix DFT (f * 2^k, f ∈ {3, 5, 15}, k ≥ 4) are also accepted; an unsupported length throws an ArgumentError. With an explicit setup::FFTSetup, or for 2D inputs, all dimensions must be powers of 2. Wraps vDSP_fft_zrop (1D) / vDSP_fft2d_zrop (2D) / vDSP_DFT_zrop_CreateSetup (1D mixed-radix).
rfft(x, setup::FFTSetup, ws::FFTWorkspace)
brfft(X, n, setup::FFTSetup, ws::FFTWorkspace)
irfft(X, n, setup::FFTSetup, ws::FFTWorkspace)Temp-buffer real FFT variants (see rfft/brfft/irfft), using the FFTWorkspace ws as scratch. Available for 1D vectors and 2D matrices; all dimensions must be powers of 2. Wraps vDSP_fft_zropt (1D) / vDSP_fft2d_zropt (2D).
rfft(x::Matrix{T}, dims::Integer, [setup, ws])
rfft!(x::Matrix{T}, dims::Integer, [setup, ws])Batched real forward FFT: transform each column (dims = 1) or row (dims = 2) of the real matrix x independently, like FFTW's rfft(x, dims). Returns the non-redundant complex coefficients ((n÷2+1) along dims). The transform length size(x, dims) must be a power of 2. rfft! operates in-place (overwriting x) using vDSP's in-place real FFT. Passing a FFTWorkspace selects the temp-buffer form. Wraps vDSP_fftm_zrop / vDSP_fftm_zropt / vDSP_fftm_zrip / vDSP_fftm_zript.
AppleAccelerate.rfft! — Function
rfft!(x::Vector{T}, [setup::FFTSetup{T}, [ws::FFTWorkspace{T}]])
rfft!(x::Matrix{T}, setup::FFTSetup{T}, ws::FFTWorkspace{T})In-place real forward FFT: vDSP transforms the packed real data within its own split-complex buffer instead of a separate output buffer, reducing internal buffer use. Returns the non-redundant complex spectrum (length n÷2+1 for a vector, size (n1÷2+1)×n2 for a matrix). The input x is overwritten (used as FFT scratch). Passing a FFTWorkspace selects the temp-buffer form. All dimensions must be powers of 2. Wraps vDSP_fft_zrip / vDSP_fft_zript (1D) / vDSP_fft2d_zript (2D).
AppleAccelerate.irfft — Function
irfft(X::Vector{Complex{T}}, n::Int, [setup::FFTSetup{T}])
irfft(X::Matrix{Complex{T}}, n1::Int, [setup::FFTSetup{T}])Compute the normalized inverse real FFT, returning a real vector of length n or a real matrix of size n1×n2 where n2 = size(X, 2). Satisfies irfft(rfft(x), size(x, 1)) ≈ x.
For a 1D input with no explicit setup, non-power-of-2 lengths supported by Apple's real-input mixed-radix DFT are also accepted (see rfft). Wraps vDSP_fft_zrop (1D) / vDSP_fft2d_zrop (2D).
AppleAccelerate.brfft — Function
brfft(X::Vector{Complex{T}}, n::Int, [setup::FFTSetup{T}])
brfft(X::Matrix{Complex{T}}, n1::Int, [setup::FFTSetup{T}])Compute the unnormalized inverse real FFT, returning a real vector of length n (X must have length n÷2+1) or a real matrix of size n1×n2 where n2 = size(X, 2) (X must have size (n1÷2+1)×n2). The result is not divided by the output length; use irfft for the normalized version.
For a 1D input with no explicit setup, non-power-of-2 lengths supported by Apple's real-input mixed-radix DFT are also accepted (see rfft). Wraps vDSP_fft_zrop (1D) / vDSP_fft2d_zrop (2D).
brfft(X::Matrix{Complex{T}}, n::Integer, dims::Integer, [setup, ws])
irfft(X::Matrix{Complex{T}}, n::Integer, dims::Integer, [setup, ws])Batched inverse real FFT along dims, the inverse of rfft(x, dims). n is the transformed (real) length along dims. brfft is unnormalized; irfft divides by n. X must have n÷2+1 entries along dims. Wraps vDSP_fftm_zrop / vDSP_fftm_zropt.
Plans: plan * x, mul!, inv
The setups returned by plan_fft/plan_rfft/plan_dft/plan_dct are vDSP's own objects: they hold twiddle tables but know neither the transform's direction nor its shape, so they can only be passed along to fft(x, setup). An FFTPlan adds that missing information, which gives every vDSP transform the plan interface Julia users know from FFTW:
| Constructor | Transform | Shapes |
|---|---|---|
fftplan, bfftplan, ifftplan | complex forward / unnormalized backward / normalized inverse | vector, matrix (2-D), matrix with dims (batched 1-D) |
rfftplan, brfftplan, irfftplan | real ↔ half spectrum | vector, matrix (2-D) |
dctplan | DCT-II / III / IV (Float32) | vector |
using LinearAlgebra: mul!
x = randn(ComplexF64, 1024)
p = AppleAccelerate.fftplan(x) # only the type and size of x are used
X = p * x # allocates the result
y = similar(x)
mul!(y, p, x) # writes into y, no allocation
@assert y == X
@assert p \ X ≈ x # same as inv(p) * X
ip = inv(p) # a normalized inverse plan; reuse it in hot loops
@assert ip * X ≈ x
z = copy(x)
mul!(z, p, z) # passing the same array transforms in place
@assert z ≈ X
pAppleAccelerate.FFTPlan{Float64}: forward FFT, 1024 ComplexF64 → 1024 ComplexF64 (vDSP_DFT_Interleaved)Plans can also be built from a type and a size, and carry the usual introspection:
julia> pr = AppleAccelerate.rfftplan(Float32, (16, 32)); # 2-D real FFT
julia> size(pr), AppleAccelerate.output_size(pr), eltype(pr)
((16, 32), (9, 32), Float32)A = randn(ComplexF32, 16, 8)
pc = AppleAccelerate.fftplan(A, 1) # FFT of every column
@assert pc * A ≈ AppleAccelerate.fft(A, 1)
@assert pc \ (pc * A) ≈ A
d = randn(Float32, 64)
pd = AppleAccelerate.dctplan(d) # DCT-II
@assert pd \ (pd * d) ≈ d # inv uses DCT-III, scaled by 2/nSupported sizes are those of the one-shot functions: 1-D complex plans take any is_supported_fft_length; 1-D real plans a power of two or a mixed-radix length accepted by Apple's real-input DFT; 2-D and batched plans powers of two; DCT plans f·2^k with f ∈ {1, 3, 5, 15}, k ≥ 4. Anything else throws an ArgumentError when the plan is constructed, never when it is applied.
No copies. vDSP's FFTs want split-complex operands, which is why the one-shot fft(x) copies x into separate real/imaginary arrays and back. A plan instead hands vDSP the interleaved Complex buffer itself as a stride-2 split-complex operand (or uses the native interleaved DFT for lengths up to 4096), so mul! performs no allocation and no packing. Up to 4096 points this is also faster than the one-shot call (about 1.4× at 1024, 2× at 960 on an M-series Mac); for larger powers of two vDSP's stride-2 access is somewhat slower than its unit-stride kernels (≈105 µs vs ≈83 µs at 16384), the price of not allocating. The exceptions are the transforms vDSP only offers without a stride argument, or with a packed layout that cannot be unpacked in place — mixed-radix lengths above the interleaved DFT's limit, mixed-radix real plans, and 2-D real plans — which go through temporaries. x and y must be contiguous; column views such as view(M, :, j) are fine.
Thread safety. A plan is immutable and owns no scratch memory, and vDSP setups are read-only while a transform executes, so one plan can be applied from many tasks concurrently. Plans share vDSP setups through the same locked caches as the one-shot API, which makes constructing a plan (or its inv) cheap.
FFTPlan does not subtype AbstractFFTs.Plan, and the constructors are deliberately not called plan_fft: AppleAccelerate defines no methods on other packages' functions for other packages' types (see Architecture). *, \, inv and mul! are extended only for FFTPlan, a type this package owns, so loading AppleAccelerate still changes nothing about FFTW's plans.
AppleAccelerate.FFTPlan — Type
FFTPlan{T,K,N,S}A reusable transform plan: a vDSP setup together with the transform kind K (:fft, :fftm, :rfft, :brfft or :dct), the direction, the input/output shape (N dimensions) and the output scaling. T is the real precision and S the vDSP setup type backing the plan.
Create one with fftplan, bfftplan, ifftplan, rfftplan, brfftplan, irfftplan or dctplan, then apply it with plan * x or mul!(y, plan, x), and invert it with inv(plan) or plan \ y. size(plan) is the input size, output_size the output size, and eltype(plan) the input element type.
Plans are immutable and hold no scratch memory, so a single plan may be applied concurrently from several tasks. The underlying vDSP setup is freed by its own finalizer once no plan references it.
AppleAccelerate.fftplan — Function
fftplan(x::StridedVecOrMat{Complex{T}}, [dims])
fftplan(Complex{T}, size, [dims])Plan a forward complex FFT for arrays shaped like x (only the element type and size of x are used; its contents are untouched). Returns an FFTPlan.
- Vector: any length accepted by
is_supported_fft_length— a power of two, orf·2^kwithf ∈ {3, 5, 15},k ≥ 3. - Matrix: a full 2-D transform; both dimensions must be powers of two.
- Matrix with
dims = 1or2: independent 1-D transforms of every column or row (power-of-two length), like FFTW'splan_fft(x, dims).
julia> using LinearAlgebra: mul!
julia> x = ComplexF64[1, 2, 3, 4];
julia> p = AppleAccelerate.fftplan(x)
AppleAccelerate.FFTPlan{Float64}: forward FFT, 4 ComplexF64 → 4 ComplexF64 (vDSP_fft)
julia> y = p * x # allocate the result
4-element Vector{ComplexF64}:
10.0 + 0.0im
-2.0 + 2.0im
-2.0 + 0.0im
-2.0 - 2.0im
julia> mul!(similar(x), p, x) == y # no allocation; `mul!(x, p, x)` transforms in place
true
julia> p \ y ≈ x # same as inv(p) * y
trueAppleAccelerate.bfftplan — Function
bfftplan(x::StridedVecOrMat{Complex{T}}, [dims])
bfftplan(Complex{T}, size, [dims])Plan an unnormalized backward (inverse) complex FFT: bfftplan(x) * (fftplan(x) * x) is x times the transform length. See fftplan for the supported shapes and ifftplan for the normalized inverse.
AppleAccelerate.ifftplan — Function
ifftplan(x::StridedVecOrMat{Complex{T}}, [dims])
ifftplan(Complex{T}, size, [dims])Plan a normalized inverse complex FFT, so that ifftplan(x) * (fftplan(x) * x) ≈ x. This is the plan inv(fftplan(x)) returns. See fftplan for the supported shapes.
AppleAccelerate.rfftplan — Function
rfftplan(x::StridedVecOrMat{T})
rfftplan(T, size)Plan a forward FFT of a real array, producing the non-redundant half spectrum in the same layout as FFTW's rfft: length n÷2+1 for a vector of length n, size (n1÷2+1)×n2 for an n1×n2 matrix.
Vector lengths may be a power of two (≥ 2) or a mixed-radix f·2^k length supported by Apple's real-input DFT (f ∈ {3, 5, 15}, k ≥ 4); matrix dimensions must be powers of two. mul! is allocation-free for power-of-two vectors; the mixed-radix and 2-D paths pack through temporaries.
inv(plan) / plan \ X give the normalized inverse (irfftplan).
AppleAccelerate.brfftplan — Function
brfftplan(X::StridedVecOrMat{Complex{T}}, n::Integer)Plan the unnormalized inverse of rfftplan: maps a half spectrum shaped like X back to a real array whose first dimension is n (size(X, 1) must be n÷2+1). The result is the original signal times the number of samples; use irfftplan for the normalized inverse.
AppleAccelerate.irfftplan — Function
irfftplan(X::StridedVecOrMat{Complex{T}}, n::Integer)Plan the normalized inverse real FFT, so that irfftplan(X, size(x, 1)) * (rfftplan(x) * x) ≈ x. See brfftplan.
AppleAccelerate.dctplan — Function
dctplan(x::StridedVector{Float32}, [dct_type = 2])
dctplan(Float32, n, [dct_type = 2])Plan an (unnormalized) discrete cosine transform of type II, III or IV (dct_type = 2, 3, 4), matching dct. The length must be f·2^k with f ∈ {1, 3, 5, 15} and k ≥ 4. vDSP provides the DCT in single precision only.
inv(plan) uses the DCT-II/DCT-III duality (and DCT-IV's self-inverse property), scaled by 2/n, so plan \ (plan * x) ≈ x.
AppleAccelerate.output_size — Function
output_size(plan::FFTPlan) -> DimsSize of the array produced by plan * x (and required of y in mul!(y, plan, x)). size(plan) is the matching input size; the two differ only for the real transforms.
Base.inv — Method
inv(plan::FFTPlan)
plan \ yinv(plan) is the plan of the inverse transform, normalized so that inv(plan) * (plan * x) ≈ x; plan \ y is inv(plan) * y. The inverse of a forward plan is the matching ifftplan/irfftplan; the inverse of an unnormalized backward plan is a forward plan scaled by 1/n. Constructing the inverse only looks up a cached vDSP setup, so it is cheap; hold on to it if you apply it in a hot loop.
LinearAlgebra.mul! — Method
mul!(y, plan::FFTPlan, x)Apply plan to x, writing the result into y and returning y. x and y must be contiguous (unit-stride) arrays of size size(plan) and output_size(plan). For the complex plans y may be x itself, which transforms in place; partially overlapping x and y are not supported.
mul! does not allocate, except for the large mixed-radix, mixed-radix real and 2-D real plans (see FFTPlan).
Unsupported lengths: an explicit fallback
vDSP cannot transform every length, and AppleAccelerate never silently switches backend. Check a length up front with is_supported_fft_length, or opt in per call site with the fallback keyword of the one-shot 1-D fft, bfft, ifft and rfft: when (and only when) vDSP cannot handle the length, fallback(x) is returned instead of an ArgumentError being thrown.
using AppleAccelerate, FFTW
x = randn(ComplexF64, 1000) # 1000 = 125·2^3 — not a vDSP length
AppleAccelerate.is_supported_fft_length(1000) # false
AppleAccelerate.fft(x) # throws ArgumentError
AppleAccelerate.fft(x; fallback = FFTW.fft) # FFTW, because you asked for it
AppleAccelerate.fft(randn(ComplexF64, 1024); fallback = FFTW.fft) # still vDSPFor plans, make the same choice when constructing them:
p = AppleAccelerate.is_supported_fft_length(length(x)) ? AppleAccelerate.fftplan(x) : FFTW.plan_fft(x)
y = p * x # both support *, mul!, inv and \Small-radix, fixed-size, and interleaved transforms
Beyond the general FFT/DFT paths, vDSP provides specialized complex-transform kernels:
- Radix-3 / radix-5 FFTs (
fftradix3,fftradix5, and their unnormalized inversesbfftradix3/bfftradix5) handle lengths3·2^kand5·2^k— including12,20,24,40, which the mixed-radixfft/dftpath cannot. - Fixed-size 16-/32-point FFTs (
fft16,fft32, inversesbfft16/bfft32) are dedicatedFloat32kernels that need no setup. - Interleaved-complex DFT (
dft_interleaved/idft_interleaved, planned withplan_dft_interleaved) operates directly onVector{Complex}buffers rather than the split-complex layout ofdft.
x = randn(ComplexF64, 12) # 12 = 3·2^2, not an fft() length
@assert AppleAccelerate.bfftradix3(AppleAccelerate.fftradix3(x)) ≈ 12 .* x
y = randn(ComplexF32, 16)
@assert AppleAccelerate.bfft16(AppleAccelerate.fft16(y)) ≈ 16 .* y
z = randn(ComplexF64, 24)
@assert AppleAccelerate.idft_interleaved(AppleAccelerate.dft_interleaved(z)) ≈ zAppleAccelerate.fftradix3 — Function
fftradix3(x::Vector{Complex{T}}) ; bfftradix3(x)
fftradix5(x::Vector{Complex{T}}) ; bfftradix5(x)Radix-3 (length 3·2^k) or radix-5 (length 5·2^k) out-of-place complex FFT. fftradix* is the forward transform; bfftradix* is the unnormalized inverse (so bfftradix3(fftradix3(x)) ≈ length(x) .* x). These cover lengths the mixed-radix DFT path cannot (e.g. 12, 20). Wraps vDSP_fft3_zop / vDSP_fft5_zop.
AppleAccelerate.fftradix5 — Function
fftradix3(x::Vector{Complex{T}}) ; bfftradix3(x)
fftradix5(x::Vector{Complex{T}}) ; bfftradix5(x)Radix-3 (length 3·2^k) or radix-5 (length 5·2^k) out-of-place complex FFT. fftradix* is the forward transform; bfftradix* is the unnormalized inverse (so bfftradix3(fftradix3(x)) ≈ length(x) .* x). These cover lengths the mixed-radix DFT path cannot (e.g. 12, 20). Wraps vDSP_fft3_zop / vDSP_fft5_zop.
AppleAccelerate.fft16 — Function
fft16(x::Vector{ComplexF32}) ; bfft16(x)
fft32(x::Vector{ComplexF32}) ; bfft32(x)Dedicated fixed-size 16-point / 32-point complex FFT (Float32 only, no setup required). fft* is the forward transform, bfft* the unnormalized inverse (bfft16(fft16(x)) ≈ 16 .* x). Wraps vDSP_FFT16_zopv / vDSP_FFT32_zopv.
AppleAccelerate.fft32 — Function
fft16(x::Vector{ComplexF32}) ; bfft16(x)
fft32(x::Vector{ComplexF32}) ; bfft32(x)Dedicated fixed-size 16-point / 32-point complex FFT (Float32 only, no setup required). fft* is the forward transform, bfft* the unnormalized inverse (bfft16(fft16(x)) ≈ 16 .* x). Wraps vDSP_FFT16_zopv / vDSP_FFT32_zopv.
AppleAccelerate.plan_dft_interleaved — Function
plan_dft_interleaved(n::Integer, direction::Integer, ::Type{T}=Float32)Create a reusable interleaved-complex DFT setup of length n and direction (DFT_FORWARD or DFT_INVERSE), for use with dft_interleaved. Throws an ArgumentError for an unsupported length. The setup is freed automatically when garbage-collected. Wraps vDSP_DFT_Interleaved_CreateSetup.
AppleAccelerate.dft_interleaved — Function
dft_interleaved(x::Vector{Complex{T}}, setup::InterleavedDFTSetup{T})
dft_interleaved(x::Vector{Complex{T}}, [direction=DFT_FORWARD])Execute an interleaved-complex DFT on x. With no setup, one is created (and cached) for the length/direction of x. Wraps vDSP_DFT_Interleaved_Execute.
AppleAccelerate.idft_interleaved — Function
idft_interleaved(x::Vector{Complex{T}})Normalized inverse interleaved-complex DFT: dft_interleaved(x, DFT_INVERSE) ./ length(x).
DFT (Complex Discrete Fourier Transform)
Wraps Apple's vDSP_DFT_zop for complex-to-complex DFT. Unlike the FFT functions, DFT supports non-power-of-2 lengths of the form f * 2^n where f ∈ {1, 3, 5, 15} and n ≥ 3. Both Float32 and Float64 are supported.
x = randn(ComplexF64, 120) # 120 = 15 * 2^3, non-power-of-2
# Forward DFT (auto-creates setup)
X = AppleAccelerate.dft(x)
# Normalized inverse DFT
x_recovered = AppleAccelerate.idft(X)
@assert x_recovered ≈ x
# Reusable setup for repeated transforms
setup_fwd = AppleAccelerate.plan_dft(120, AppleAccelerate.DFT_FORWARD, Float64)
setup_inv = AppleAccelerate.plan_dft(120, AppleAccelerate.DFT_INVERSE, Float64)
X = AppleAccelerate.dft(x, setup_fwd)
x_recovered = AppleAccelerate.idft(X, setup_inv)AppleAccelerate.plan_dft — Function
plan_dft(length::Int, direction::Int, ::Type{T}=Float32; previous=C_NULL) where TCreate a DFT setup for complex-to-complex DFT of the given length and direction (DFT_FORWARD or DFT_INVERSE). Length must be f * 2^k where f ∈ {3, 5, 15} and k ≥ 3 (a power of two is also accepted). Optionally pass a previous setup to share underlying data. Wraps vDSP_DFT_zop_CreateSetup.
Returns: DFTSetup{T}
AppleAccelerate.dft — Function
dft(Ir::Vector{T}, Ii::Vector{T}, setup::DFTSetup{T})Execute the complex DFT defined by setup on split-complex input (Ir, Ii). Wraps vDSP_DFT_Execute.
Returns: (Or, Oi) — real and imaginary parts of the output.
dft(X::Vector{Complex{T}}, setup::DFTSetup{T})Execute the complex DFT on an interleaved complex vector. Wraps vDSP_DFT_Execute.
Returns: Vector{Complex{T}}
dft(X::Vector{Complex{T}}, direction::Int) where T
dft(X::Vector{Complex{T}}) where TCompute the DFT of X, auto-creating a setup. Default direction is forward. Wraps vDSP_DFT_Execute.
Returns: Vector{Complex{T}}
AppleAccelerate.idft — Function
idft(X::Vector{Complex{T}}, setup::DFTSetup{T})
idft(X::Vector{Complex{T}})Compute the normalized inverse DFT of X. The setup must have been created with DFT_INVERSE direction. Returns dft(X, setup) ./ length(X). Wraps vDSP_DFT_Execute.
Returns: Vector{Complex{T}}
DCT (Discrete Cosine Transform)
Wraps Apple's vDSP DCT functions. Float32 only. Supports DCT types II, III, and IV.
| Function | Description |
|---|---|
plan_dct | Create a DCT setup object |
dct | Compute the Discrete Cosine Transform |
idct | Compute the inverse Discrete Cosine Transform |
plan_destroy | Destroy a DCT setup object |
AppleAccelerate.plan_dct — Function
plan_dct(length, dct_type, [previous])Create a DCT setup object. dct_type must be 2, 3, or 4 (Type II, III, IV). Length must be f * 2^n where f ∈ {1,3,5,15} and n ≥ 4. Wraps vDSP_DCT_CreateSetup.
AppleAccelerate.dct — Function
dct(X::Vector{Float32}, setup::DFTSetup)
dct(X::Vector{Float32}, [dct_type=2])Compute the Discrete Cosine Transform of X. Wraps vDSP_DCT_Execute.
Apple's Accelerate framework provides only a single-precision DCT (vDSP_DCT_CreateSetup/vDSP_DCT_Execute); there is no Float64 variant (no vDSP_DCT_CreateSetupD/vDSP_DCT_ExecuteD exists in vDSP). Calling dct or idct with a Vector{Float64} therefore throws an ArgumentError. Convert the input to Float32, or use FFTW.jl for a double-precision DCT.
AppleAccelerate.idct — Function
idct(X::Vector{Float32})Compute the inverse of the (unnormalized) DCT-II computed by dct, so that idct(dct(x)) ≈ x. Uses the DCT-II/DCT-III duality: vDSP's DCT-III applied to a DCT-II spectrum reproduces the input scaled by N/2, so this returns dct(X, 3) .* (2/N). Wraps vDSP_DCT_Execute.
Not available for Float64 inputs: see the note in dct.
AppleAccelerate.plan_destroy — Function
plan_destroy(setup::DFTSetup)Destroy a DCT/DFT setup object, freeing its resources. Wraps vDSP_DFT_DestroySetup.
Allocation-Free Fixed-Size FFTs
The mutating forms of the 16-/32-point transforms use vDSP's interleaved-complex kernels, which read and write the memory of a Vector{ComplexF32} directly — no split-complex repacking and no allocation. Use them in tight per-block loops.
| Function | Description |
|---|---|
fft16! / fft32! | Forward transform, in place (f!(x)) or into out (f!(out, x)) |
bfft16! / bfft32! | Unnormalized inverse, same two forms |
using AppleAccelerate
x = randn(ComplexF32, 16); x0 = copy(x)
AppleAccelerate.fft16!(x) # in place
AppleAccelerate.bfft16!(x) # back again, scaled by 16
x ≈ 16 .* x0trueAppleAccelerate.fft16! — Function
fft16!(x) ; fft16!(out, x) ; bfft16!(x) ; bfft16!(out, x)
fft32!(x) ; fft32!(out, x) ; bfft32!(x) ; bfft32!(out, x)Allocation-free forms of the fixed-size fft16 / fft32 transforms. x and out are length-16 (resp. 32) ComplexF32 vectors; the one-argument form transforms x in place. fft*! is the forward transform and bfft*! the unnormalized inverse (bfft16!(fft16!(x)) ≈ 16 .* x).
These use vDSP's interleaved-complex kernels, which operate directly on the memory of a Vector{ComplexF32} with no split-complex repacking. Unit-stride views are accepted provided they are 16-byte aligned (any view starting at an even element offset of a Vector{ComplexF32} is); out and x may be the same array but must not otherwise overlap.
Wraps vDSP_FFT16_copv / vDSP_FFT32_copv.
AppleAccelerate.bfft16! — Function
fft16!(x) ; fft16!(out, x) ; bfft16!(x) ; bfft16!(out, x)
fft32!(x) ; fft32!(out, x) ; bfft32!(x) ; bfft32!(out, x)Allocation-free forms of the fixed-size fft16 / fft32 transforms. x and out are length-16 (resp. 32) ComplexF32 vectors; the one-argument form transforms x in place. fft*! is the forward transform and bfft*! the unnormalized inverse (bfft16!(fft16!(x)) ≈ 16 .* x).
These use vDSP's interleaved-complex kernels, which operate directly on the memory of a Vector{ComplexF32} with no split-complex repacking. Unit-stride views are accepted provided they are 16-byte aligned (any view starting at an even element offset of a Vector{ComplexF32} is); out and x may be the same array but must not otherwise overlap.
Wraps vDSP_FFT16_copv / vDSP_FFT32_copv.
AppleAccelerate.fft32! — Function
fft16!(x) ; fft16!(out, x) ; bfft16!(x) ; bfft16!(out, x)
fft32!(x) ; fft32!(out, x) ; bfft32!(x) ; bfft32!(out, x)Allocation-free forms of the fixed-size fft16 / fft32 transforms. x and out are length-16 (resp. 32) ComplexF32 vectors; the one-argument form transforms x in place. fft*! is the forward transform and bfft*! the unnormalized inverse (bfft16!(fft16!(x)) ≈ 16 .* x).
These use vDSP's interleaved-complex kernels, which operate directly on the memory of a Vector{ComplexF32} with no split-complex repacking. Unit-stride views are accepted provided they are 16-byte aligned (any view starting at an even element offset of a Vector{ComplexF32} is); out and x may be the same array but must not otherwise overlap.
Wraps vDSP_FFT16_copv / vDSP_FFT32_copv.
AppleAccelerate.bfft32! — Function
fft16!(x) ; fft16!(out, x) ; bfft16!(x) ; bfft16!(out, x)
fft32!(x) ; fft32!(out, x) ; bfft32!(x) ; bfft32!(out, x)Allocation-free forms of the fixed-size fft16 / fft32 transforms. x and out are length-16 (resp. 32) ComplexF32 vectors; the one-argument form transforms x in place. fft*! is the forward transform and bfft*! the unnormalized inverse (bfft16!(fft16!(x)) ≈ 16 .* x).
These use vDSP's interleaved-complex kernels, which operate directly on the memory of a Vector{ComplexF32} with no split-complex repacking. Unit-stride views are accepted provided they are 16-byte aligned (any view starting at an even element offset of a Vector{ComplexF32} is); out and x may be the same array but must not otherwise overlap.
Wraps vDSP_FFT16_copv / vDSP_FFT32_copv.