An experimental First-Order Ambisonics (FOA) Sound Event Localization and Detection (SELD) system for stateful streaming inference and Core ML deployment.
This project is developed as part of my Master's thesis, aiming to push SELD from conventional offline/window-based evaluation toward causal, low-latency, online inference on edge devices.
The system maintains two independent inference paths:
- Stateless SELD — conventional sliding-window inference.
- Stateful causal SELD — incremental inference with persistent
MLStateand cross-chunk computation reuse.
This is a research prototype. Contributions, device testing, model integrations, and discussions are welcome.
The default stateful_cstformer_mel64 configuration is:
| Item | Configuration |
|---|---|
| FOA input | 4 channels, 24 kHz, 16-bit PCM |
| STFT | n_fft=1024, win_length=960, hop=480 |
| Frame interval | 20 ms, center=False |
| Features | 4-channel log-Mel + 3-channel FOA intensity vector |
| Feature tensor | [1, 7, T, 64] |
| Core ML micro-chunk | [1, 7, 5, 64] = 100 ms |
| Output | Multi-ACCDOA [1, 1, 117] |
streamChunkFrames supports 5/10/25/50/100 frames. Application chunks are processed as fixed 5-frame Core ML micro-chunks while reusing the same MLState.
FOA Audio
→ Streaming Features
→ [1, 7, T, 64]
→ 100 ms Micro-chunks
→ Stateful Core ML
→ Multi-ACCDOA
Changing the application chunk size affects scheduling and batching overhead, but does not change model weights or architecture.
The stateless path uses fixed sliding windows with configurable inference strides of 100/200/300/500/1000 ms.
AUHAL / AVAudioEngine / WAV
↓
FOAAudioFrame
↓
FOARealtimeSELDProcessor
↓
MelStreamingFOAFeatureExtractor
↓
FOAFeatureTensor
↓
CoreMLFOASELDPredictor + MLState
↓
Multi-ACCDOA Decoder
↓
3D Visualization / CSV / Profiling
Main components:
| Component | Location |
|---|---|
| Model configuration | Resources/Models/Models.json |
| Mel configuration | Resources/Features/Mel64Streaming.json |
| Streaming features | Shared/Features/MelStreamingFOAFeatureExtractor.swift |
| Streaming scheduler | Shared/SELD/FOARealtimeSELDProcessor.swift |
Core ML / MLState |
Shared/ModelDeployment/CoreMLFOASELDPredictor.swift |
| Audio input | Shared/Audio/ |
| Dataset evaluation | Shared/SELD/FOADatasetEvaluation.swift |
| Performance profiling | Shared/Performance/StageProfiler.swift |
The application supports:
- Real-time and offline FOA WAV inference
- Stateful and stateless model comparison
- Configurable inference chunk/stride
- CPU / GPU backend selection
- 3D SELD visualization
- DCASE-format prediction export
- Stage-level latency, RTF, CPU and memory profiling
- Full-dataset evaluation
Each evaluation file resets feature history and MLState to reproduce streaming cold-start conditions.
Results are exported as:
results/<runName>/
├── predictions/*.csv
└── run_report.json
The deployment pipeline verifies numerical consistency between Python training features and the Swift/vDSP implementation.
Current tolerance:
MAE < 2e-4
Max Error < 5e-3
Run the tests with:
xcodebuild test \
-project 'FOASELD System.xcodeproj' \
-scheme 'FOASELD System' \
-destination 'platform=macOS,arch=arm64' \
-only-testing:'FOASELD SystemTests' \
CODE_SIGNING_ALLOWED=NOThis Master's thesis explores a practical question:
How can SELD move from offline evaluation toward continuous, causal, low-latency inference on real edge devices?
The project focuses on:
- Stateful and causal SELD
- Streaming feature extraction
- Cross-chunk state reuse
- Short-context and incremental inference
- Cold-start behavior
- Core ML deployment
- End-to-end latency, RTF, CPU and memory evaluation
The goal is to connect algorithmic efficiency with measurable on-device performance, while keeping conventional stateless models as reproducible baselines.
Contributions and independent testing are welcome, especially for:
- Different Apple devices
- Alternative SELD architectures
- Core ML stateful models
- Additional FOA datasets
- Streaming and latency optimization
- CPU / GPU / ANE profiling
- Reproducibility and feature-parity testing
Feel free to open an issue, report experimental results, or submit a pull request.
Goal: make SELD not only accurate offline, but practical for continuous low-latency perception on real devices.