This project benchmarks the performance of 2D convolutions on Apple Silicon using Metal Performance Shaders Graph (MPSGraph). It measures the compute capacity in TOPS (Trillions of Operations Per Second) for both the GPU and the Apple Neural Engine (ANE).
To compile the project, ensure you have clang and the necessary frameworks (Foundation, Metal, MetalPerformanceShadersGraph) installed (standard on macOS with Xcode Command Line Tools).
makeTo build just the Swift version:
make measure_conv_swiftTo clean the build artifacts:
make cleanTo build for iOS, you need to use the xcrun command to target the iPhone SDK and sign the binary.
# Compile for iOS (arm64)
xcrun -sdk iphoneos clang -fobjc-arc -O3 -framework Foundation -framework Metal -framework MetalPerformanceShadersGraph measure_conv_universal.m -o measure_conv_ios
# Sign the binary (replace 'Apple Development' with your identity)
codesign -s "Apple Development" measure_conv_iosNote: Running a standalone binary on a non-jailbroken iPhone is restricted. The easiest way to run this on a device is to wrap it in an iOS App.
- Open Xcode and create a new iOS App (Objective-C).
- Delete the following default files:
AppDelegate.h/m,SceneDelegate.h/m, andViewController.h/m. - Replace the contents of
main.mwith the code frommeasure_conv_gui.m. - Info.plist (Scene Manifest):
- In the Info tab, find Application Scene Manifest.
- Delete that entire row (to prevent the app from trying to use a SceneDelegate).
- Add
MetalandMetalPerformanceShadersGraphto the Frameworks, Libraries, and Embedded Content. - Run the app on your connected iPhone.
- Follow the steps above but use
measure_conv_universal.minstead. - Check the Xcode Console for the output.
./measure_convmeasure_ane_pmu provides deep physical hardware profiling for Apple Neural Engine via _ANEClient and Apple PMU registers (com.apple.ane.hardware-counters), comparing FP16, INT8, QDQ, and GPU baselines while exporting self-contained .mpsgraphpackage bundles.
Note / Attribution:
measure_ane_pmuis based on and adapted from ane_pmu_profiler, reusing its low-level Apple Neural Engine silicon PMU profiling, private_ANEClienttelemetry interfaces, and hardware register mapping.For an in-depth microarchitectural analysis explaining how these counters behave under different tensor dimensions, see the Guide to Interpreting measure_ane_pmu Numbers.
# Build and codesign with PMU entitlements
make measure_ane_pmu
# Run full benchmark across FP16, INT8, and QDQ with 20 chained layers
./measure_ane_pmu --layers 20 --iterations 10
# Export .mpsgraphpackage models to custom directory
./measure_ane_pmu --variant all --save-package ./packages
# Profile a specific variant without GPU
./measure_ane_pmu --variant int8 --layers 20 --iterations 20 --no-gpu| Flag | Description | Default |
|---|---|---|
--variant <type> |
Precision variant (all, fp16, int8, qdq) |
all |
--batch <B> |
Batch dimension | 1 |
--size <H> |
Spatial height and width ( |
256 |
--channels <C> |
Input and output channel depth ( |
128 |
--layers <L> |
Number of chained convolution layers | 20 |
--iterations <N> |
Number of benchmark passes | 20 |
--save-package <dir> |
Serialize .mpsgraphpackage bundles |
./packages |
--no-gpu |
Skip Metal GPU comparison | Disabled |
--no-pmu |
Skip physical silicon PMU hardware profiling | Disabled |
When --save-package is enabled, measure_ane_pmu serializes each variant into a self-contained bundle containing:
manifest.plist& compiled graph bytecode (original_model_0.mpsgraph,specialized_model_1.mpsgraph)resources.binwith constant weightsane_bundle/: Contains the low-level ANECIR bitcode (*.bc.mlir) and compiler options (compiler_options_*.plist) emitted by the MPSGraph compiler, allowing direct replay with_ANEClient.
Scaling channel depth ($C=256$) and layer count ($L=50$) maximizes arithmetic intensity, completely amortizes command queue dispatch, and pushes the physical silicon ALU arrays to their architectural limits (3.865 TOPs / 1.933 Trillion MACs per pass):
| Variant | Precision | Latency | Speed (TOPS) | Output Stalls ([15]) |
DMA Traffic ([17]) |
Throughput / Core ([10]) |
Total Chip Throughput | Peak Saturation |
|---|---|---|---|---|---|---|---|---|
| ANE FP16 | Float16 | 205.48 ms | 18.81 | 5,346,498,518 | 350.10 MB | 248.9 MACs/cyc/core | 3,982.2 MACs/cycle | 97.22% (of 256 peak) |
| ANE INT8 | Int8 | 101.67 ms | 38.02 🏆 | 1,914,275,704 | 173.11 MB | 503.4 MACs/cyc/core | 8,054.1 MACs/cycle | 98.32% (of 512 peak) |
🏆 Record Peak Reached: Native INT8 achieves 38.02 TOPS, hitting 100.05% of Apple's advertised 38 TOPS ceiling on Apple M4 Pro silicon, with each of the 16 cores executing 503.4 MACs / cycle out of the 512 physical hardware maximum.
Workload: 20 chained 3×3 Conv layers on [1, 128, 256, 256] tensor (386.55 GOPs / 0.3865 TOPs per pass)
| Variant | Device | Latency | Speed (TOPS) | Compute Cycles* | Output Stalls | Planar Cycles (L2PE) | DMA Traffic | Throughput / Core ([10]) |
|---|---|---|---|---|---|---|---|---|
| GPU FP16 | Metal GPU | 39.13 ms | 9.88 | — | — | — | — | — |
| ANE FP16 | Physical ANE | 20.61 ms | 18.76 | 157,200* | 503,317,431 | 10,494,464 | 34.99 MB | 248.2 MACs/cyc/core (97.0%) |
| ANE INT8 | Physical ANE | 10.78 ms | 35.87 | 105,045,873 | 233,509,350 | 5,247,232 | 18.35 MB | 472.0 MACs/cyc/core (92.2%) |
| ANE QDQ | Physical ANE | 20.78 ms | 18.60 | 5,010,176 | 627,454,283 | 5,249,536 | 35.27 MB | 246.1 MACs/cyc/core (96.1%) |
*Note on Compute Cycles: Under memory-bound workloads (16 MB intermediate feature maps spilling on-chip L2 SRAM to DRAM), kANE_NE_COMPUTE_CYCLES is clock-gated OFF during the 503M output writeback stall cycles. In conv_fp16, back-to-back convolutions keep output writeback queues continuously saturated. In conv_int8, intermediate requantization casts on the Planar Engine create distinct computational phases, allowing the convolution engine to log 105M unstalled cycles. Total pipeline cycles (Compute + Stalls) account for 100% of runtime across all variants.
When intermediate feature maps (~1.0 MB) fit entirely within on-chip L2 SRAM (~4–8 MB), output stalls collapse by >16× and unstalled compute cycles become directly visible:
| Variant (64×64, L=10) | Latency | Speed (TOPS) | Compute Cycles ([13]) |
Output Stalls ([15]) |
DMA Traffic ([17]) |
Throughput / Core ([10]) |
Total Chip Throughput |
|---|---|---|---|---|---|---|---|
| ANE FP16 | 1.15 ms | 10.50 | 463,416 | 15,587,386 | 1.80 MB | 156.5 MACs/cyc/core | 2,504.0 MACs/cyc (61.1% peak) |
| ANE INT8 | 0.77 ms | 15.72 | 484,910 | 7,547,107 | 1.23 MB | 212.2 MACs/cyc/core | 3,395.2 MACs/cyc (41.5% peak) |
Key Architectural Observations from Silicon PMU:
- Peak INT8 Realization: Native INT8 reaches 38.02 TOPS on Apple M4 Pro (503.4 MACs/cycle/core), fully saturating the 38 TOPS hardware specification (100.05%).
- Peak FP16 Realization: Native FP16 reaches 18.81 TOPS (248.9 MACs/cycle/core out of 256 physical limit), operating at 97.22% ALU saturation.
- Integer Scaling & DMA Reduction: INT8 doubles throughput over FP16 and cuts Unified Memory DMA traffic directly in half.
- Output Backpressure Stalls: Because large feature maps exceed on-chip L2 SRAM (~4–8 MB), large spatial maps incur output backpressure to DRAM (
kANE_NE_OUTPUT_STALL_CYCLES). Halving tensor size in INT8 cuts output stalls by >2.1× (503M → 233M cycles).- L2 SRAM Fitting: Reducing spatial dimensions to fit inside L2 SRAM (
$H=64, W=64$ ) collapses output writeback stalls by >16× (15.5M cycles). Note that dividing Total MACs byCOMPUTE_CYCLES([13]) yields an inflated ratio because[13]is gated during stalls; the physically bounded metric is Throughput per Nominal Cycle ([10]).- QDQ Execution: In QDQ (
dequantize -> conv -> quantize), the internal convolution arithmetic executes in FP16 precision, matching FP16 throughput (~18.60 TOPS) and FP16 DMA footprint (~35.27 MB).(For an exhaustive breakdown of each register, see
How_to_Interpret_measure_ane_pmu_Numbers.md.)
| Model | Device | Precision | Latency (Avg) | Speed (TOPS) |
|---|---|---|---|---|
| Mac Mini M4 Pro | GPU | FP16 | 39.05 ms | 9.90 |
| Mac Mini M4 Pro | ANE | FP16 | 20.90 ms | 18.50 |
| Mac Mini M4 Pro | ANE | INT8 | 10.76 ms | 35.91 |
| Mac Mini M4 Pro | ANE | DQ->FP16->Q | 20.65 ms | 18.72 |
| MacBook Pro M1 | GPU | FP16 | 129.10 ms | 2.99 |
| MacBook Pro M1 | ANE | FP16 | 35.70 ms | 10.83 |
| MacBook Pro M1 | ANE | INT8 | 34.02 ms | 11.36 |
| MacBook Pro M1 | ANE | DQ->FP16->Q | 33.99 ms | 11.37 |
| iPhone 16 Pro | GPU | FP16 | 149.35 ms | 2.59 |
| iPhone 16 Pro | ANE | FP16 | 12.81 ms | 30.18 |
| iPhone 16 Pro | ANE | INT8 | 7.98 ms | 48.46 |
| iPhone 16 Pro | ANE | DQ->FP16->Q | 6.41 ms | 60.29 |
| iPhone 17 Pro | GPU | FP16 | 57.02 ms | 6.78 |
| iPhone 17 Pro | ANE | FP16 | 8.70 ms | 44.41 |
| iPhone 17 Pro | ANE | INT8 | 7.85 ms | 49.27 |
| iPhone 17 Pro | ANE | DQ->FP16->Q | 6.12 ms | 63.20 |
Note: Results may vary slightly depending on system load and thermal state. Note: GPU INT8 convolution is not supported by MPSGraph on this device/configuration.