Daily Perf Improver: Research and Plan
Performance Testing Infrastructure
Current Testing Setup
- Build System: FAKE-based build using build/build.fs with .NET 8.0
- Test Framework: NUnit with FsUnit for tests, GitHubActionsTestLogger for CI
- CI/CD: GitHub Actions on Windows and Ubuntu (pull-requests.yml, push-master.yml)
- Build Commands:
build.sh -t Build / build.cmd -t Build
dotnet run --project build/build.fsproj -- -t RunTests
dotnet run --project build/build.fsproj -- -t All
Available Tools for Performance Work
- Formatting: Fantomas for code formatting
- Documentation: FsDocs for documentation generation
- No BenchmarkDotNet: Currently no micro-benchmarking framework detected
- Basic Performance Measurement: Some Stopwatch usage found in tests (HtmlParser.fs, IO.fs)
FSharp.Data Architecture and Performance-Critical Areas
Core Components
-
JSON Processing (FSharp.Data.Json.Core/)
- JsonValue.fs: Core JSON representation and parsing
- JsonConversions.fs: Type conversions
- JsonInference.fs: Type inference engine
- JsonRuntime.fs: Runtime operations
-
CSV Processing (FSharp.Data.Csv.Core/)
- CsvFile.fs: Core CSV file handling
- CsvInference.fs: Type inference
- CsvRuntime.fs: Runtime operations
-
HTML Processing (FSharp.Data.Html.Core/)
- HtmlParser.fs: HTML parsing engine
- HtmlCssSelectors.fs: CSS selector engine
- HtmlCharRefs.fs: Character reference handling
-
XML Processing (FSharp.Data.Xml.Core/)
- XmlInference.fs: XML type inference
- XsdInference.fs: XSD schema inference
-
Common Runtime (FSharp.Data.Runtime.Utilities/)
- StructuralInference.fs: Core type inference algorithms
- TextConversions.fs: Text parsing and conversion utilities
- IO.fs: File and URI resolution with caching
- Caching.fs: Caching mechanisms
-
HTTP (FSharp.Data.Http/)
- Http.fs: HTTP client functionality
-
Design-Time Components (FSharp.Data.DesignTime/)
- Type provider implementations for all formats
Historical Performance Improvements
From release notes analysis, significant performance work has been done:
- JSON: 20% parsing performance improvement (v2.3.0-beta2), improved JsonValue.Parse() and ToString() performance
- HTML: Improved CDATA parsing performance (v3.0.0-beta3)
- CSV: Support for large CSV files, streaming mode implementation
- General: Multiple "performance improvements" mentioned across versions
- Memory: Type provider design time component memory leak fixes and performance improvements
- Number Parsing: Improved performance of number and DateTime parsing
Typical Workloads and Bottlenecks
Primary Use Cases:
- Type Provider Design-Time: Schema inference from sample data (JSON, XML, CSV, HTML)
- Runtime Data Processing: Parsing and processing data files
- HTTP Data Access: Fetching and processing remote data sources
- Large File Handling: Processing big CSV files, large JSON documents
Likely Performance Bottlenecks:
- Type Inference: Complex structural inference algorithms in StructuralInference.fs
- String Processing: Heavy text parsing and conversion operations
- Memory Allocation: Creating intermediate objects during parsing
- I/O Operations: File reading and HTTP requests
- Reflection: Type provider code generation and runtime operations
Performance Goals and Priorities
Round 1 (Low-hanging fruit):
- Add BenchmarkDotNet for micro-benchmarking infrastructure
- Profile and optimize JSON parsing hot paths
- Improve string allocation patterns in parsers
- Optimize common type conversion operations
Round 2 (Medium complexity):
- Optimize structural inference algorithms
- Improve CSV streaming performance for large files
- Enhance HTML parser efficiency
- Reduce memory allocations in type providers
Round 3 (Advanced):
- Implement vectorization for numeric parsing where applicable
- Optimize HTTP client performance
- Improve caching mechanisms
- Advanced parser optimizations (JSON, XML, HTML)
Measurement Strategy
Current Constraints:
- No dedicated performance testing infrastructure
- Tests run in virtualized GitHub Actions environment (timing may vary)
- Need to establish baseline measurements
Proposed Approach:
- Add BenchmarkDotNet to test dependencies
- Create performance test suite covering key scenarios:
- JSON parsing of various document sizes
- CSV processing of large files
- HTML parsing of complex documents
- Type inference on representative samples
- Focus on relative improvements rather than absolute timing
- Use memory profiling to identify allocation hot spots
Environment Setup Steps
Prerequisites:
- .NET 8.0 SDK (specified in global.json)
- Restore tools:
dotnet tool restore
- Restore packages:
dotnet paket restore
Performance Development Workflow:
- Build:
dotnet run --project build/build.fsproj -- -t Build
- Run tests:
dotnet run --project build/build.fsproj -- -t RunTests
- Format code:
dotnet run --project build/build.fsproj -- -t Format
- [Need to add] Run benchmarks: TBD after BenchmarkDotNet integration
Repository Maintainer Focus Areas
Based on historical release notes, maintainers prioritize:
- Correctness: Maintaining spec compliance (JSON, CSV standards)
- Memory Management: Preventing leaks in type providers
- Large File Support: Streaming and efficient processing
- API Stability: Maintaining backward compatibility
- Cross-platform: Windows/Linux support
Next Steps
- Set up BenchmarkDotNet infrastructure for reliable micro-benchmarking
- Profile existing hot paths using sample data
- Create baseline performance measurements
- Identify specific optimization targets based on profiling results
- Implement targeted optimizations with before/after measurements
Testing Commands
# Build and test
dotnet run --project build/build.fsproj -- -t All
# Run tests only
dotnet run --project build/build.fsproj -- -t RunTests
# Format code
dotnet run --project build/build.fsproj -- -t Format
# Check formatting
dotnet run --project build/build.fsproj -- -t CheckFormat
AI-generated content by Daily Perf Improver may contain mistakes.
Daily Perf Improver: Research and Plan
Performance Testing Infrastructure
Current Testing Setup
build.sh -t Build/build.cmd -t Builddotnet run --project build/build.fsproj -- -t RunTestsdotnet run --project build/build.fsproj -- -t AllAvailable Tools for Performance Work
FSharp.Data Architecture and Performance-Critical Areas
Core Components
JSON Processing (
FSharp.Data.Json.Core/)CSV Processing (
FSharp.Data.Csv.Core/)HTML Processing (
FSharp.Data.Html.Core/)XML Processing (
FSharp.Data.Xml.Core/)Common Runtime (
FSharp.Data.Runtime.Utilities/)HTTP (
FSharp.Data.Http/)Design-Time Components (
FSharp.Data.DesignTime/)Historical Performance Improvements
From release notes analysis, significant performance work has been done:
Typical Workloads and Bottlenecks
Primary Use Cases:
Likely Performance Bottlenecks:
Performance Goals and Priorities
Round 1 (Low-hanging fruit):
Round 2 (Medium complexity):
Round 3 (Advanced):
Measurement Strategy
Current Constraints:
Proposed Approach:
Environment Setup Steps
Prerequisites:
dotnet tool restoredotnet paket restorePerformance Development Workflow:
dotnet run --project build/build.fsproj -- -t Builddotnet run --project build/build.fsproj -- -t RunTestsdotnet run --project build/build.fsproj -- -t FormatRepository Maintainer Focus Areas
Based on historical release notes, maintainers prioritize:
Next Steps
Testing Commands