Skip to content

Repository files navigation

FOA SELD: Stateful Streaming & On-Device Evaluation

An experimental First-Order Ambisonics (FOA) Sound Event Localization and Detection (SELD) system for stateful streaming inference and Core ML deployment.

This project is developed as part of my Master's thesis, aiming to push SELD from conventional offline/window-based evaluation toward causal, low-latency, online inference on edge devices.

The system maintains two independent inference paths:

  • Stateless SELD — conventional sliding-window inference.
  • Stateful causal SELD — incremental inference with persistent MLState and cross-chunk computation reuse.

This is a research prototype. Contributions, device testing, model integrations, and discussions are welcome.


1. Stateful Streaming Configuration

The default stateful_cstformer_mel64 configuration is:

Item Configuration
FOA input 4 channels, 24 kHz, 16-bit PCM
STFT n_fft=1024, win_length=960, hop=480
Frame interval 20 ms, center=False
Features 4-channel log-Mel + 3-channel FOA intensity vector
Feature tensor [1, 7, T, 64]
Core ML micro-chunk [1, 7, 5, 64] = 100 ms
Output Multi-ACCDOA [1, 1, 117]

streamChunkFrames supports 5/10/25/50/100 frames. Application chunks are processed as fixed 5-frame Core ML micro-chunks while reusing the same MLState.

FOA Audio
  → Streaming Features
  → [1, 7, T, 64]
  → 100 ms Micro-chunks
  → Stateful Core ML
  → Multi-ACCDOA

Changing the application chunk size affects scheduling and batching overhead, but does not change model weights or architecture.

The stateless path uses fixed sliding windows with configurable inference strides of 100/200/300/500/1000 ms.


2. System Architecture

AUHAL / AVAudioEngine / WAV
        ↓
FOAAudioFrame
        ↓
FOARealtimeSELDProcessor
        ↓
MelStreamingFOAFeatureExtractor
        ↓
FOAFeatureTensor
        ↓
CoreMLFOASELDPredictor + MLState
        ↓
Multi-ACCDOA Decoder
        ↓
3D Visualization / CSV / Profiling

Main components:

Component Location
Model configuration Resources/Models/Models.json
Mel configuration Resources/Features/Mel64Streaming.json
Streaming features Shared/Features/MelStreamingFOAFeatureExtractor.swift
Streaming scheduler Shared/SELD/FOARealtimeSELDProcessor.swift
Core ML / MLState Shared/ModelDeployment/CoreMLFOASELDPredictor.swift
Audio input Shared/Audio/
Dataset evaluation Shared/SELD/FOADatasetEvaluation.swift
Performance profiling Shared/Performance/StageProfiler.swift

3. Evaluation

The application supports:

  • Real-time and offline FOA WAV inference
  • Stateful and stateless model comparison
  • Configurable inference chunk/stride
  • CPU / GPU backend selection
  • 3D SELD visualization
  • DCASE-format prediction export
  • Stage-level latency, RTF, CPU and memory profiling
  • Full-dataset evaluation

Each evaluation file resets feature history and MLState to reproduce streaming cold-start conditions.

Results are exported as:

results/<runName>/
├── predictions/*.csv
└── run_report.json

4. Swift / Python Feature Parity

The deployment pipeline verifies numerical consistency between Python training features and the Swift/vDSP implementation.

Current tolerance:

MAE       < 2e-4
Max Error < 5e-3

Run the tests with:

xcodebuild test \
  -project 'FOASELD System.xcodeproj' \
  -scheme 'FOASELD System' \
  -destination 'platform=macOS,arch=arm64' \
  -only-testing:'FOASELD SystemTests' \
  CODE_SIGNING_ALLOWED=NO

5. Research Motivation

This Master's thesis explores a practical question:

How can SELD move from offline evaluation toward continuous, causal, low-latency inference on real edge devices?

The project focuses on:

  • Stateful and causal SELD
  • Streaming feature extraction
  • Cross-chunk state reuse
  • Short-context and incremental inference
  • Cold-start behavior
  • Core ML deployment
  • End-to-end latency, RTF, CPU and memory evaluation

The goal is to connect algorithmic efficiency with measurable on-device performance, while keeping conventional stateless models as reproducible baselines.


Contributing

Contributions and independent testing are welcome, especially for:

  • Different Apple devices
  • Alternative SELD architectures
  • Core ML stateful models
  • Additional FOA datasets
  • Streaming and latency optimization
  • CPU / GPU / ANE profiling
  • Reproducibility and feature-parity testing

Feel free to open an issue, report experimental results, or submit a pull request.

Goal: make SELD not only accurate offline, but practical for continuous low-latency perception on real devices.

About

First Order Ambisonics multi-channel audio based Sound Event Localization and Detection pipeline on iOS/MacOS operation system.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages