Skip to content

Fix Windows/MSVC handling of large files - #2748

Open
aydgelee wants to merge 6 commits into
WebAssembly:mainfrom
aydgelee:fix/windows-large-file-io
Open

aydgelee wants to merge 6 commits into
WebAssembly:mainfrom
aydgelee:fix/windows-large-file-io

Conversation

@aydgelee

@aydgelee aydgelee commented May 28, 2026 •

Copy link
Copy Markdown

Problem

Windows x64 builds can fail to read WAT files larger than 2GB (value too large / unable to read file). A 64-bit process does not make MSVC's long, fseek, ftell, or the default _stat file-size fields 64-bit.

This is motivated by actual use of input files larger than 4GB, which exposed the bug. The earlier documented 2,978,116,509-byte input is a validation sample, not the largest real-world input. The patch repairs file IO; it does not redesign the parser to stream inputs or reduce its whole-file memory requirements.

Changes

  • Redirect stat, fseek, and ftell to _stat64, _fseeki64, and _ftelli64 in common.cc under MSVC. Infer the file-size type from ftell and reject sizes that cannot fit in size_t before allocating.
  • Redirect fseek to _fseeki64 in stream.cc, using the same seek call across platforms. Remove the extra SeekFile helper. On the non-MSVC path, check the actual seek offset (at) against LONG_MAX before passing it to fseek.
  • Remove the additional Windows-only chunked read/write loops. Microsoft's published UCRT source package already bounds low-level transfers in stdio/fread.cpp and stdio/fwrite.cpp. The low-level _read limit does not establish a limit on an entire fread call. This revision keeps the regular fread / fwrite paths shared across platforms.
  • Merge current upstream main and adapt to the ByteSpan API. The 8MB stack reserve in wabt_executable is already upstream via #2815, so this PR no longer contains a separate stack-size change.

Validation

  • Follow-up commit 2d6c0725 passed the complete Windows x64 CI job, including 138 unit tests, C API tests, and the full test suite. The >4GB seek/backpatch test ran successfully in 5ms and the revised 2GB+ whole-file read test in 1131ms; neither was skipped.
  • All 17 GitHub checks passed on the previous revision cbe0a753 (CI run). Checks for the follow-up revision will be reported separately.
  • Local macOS Debug builds of wat2wasm, wasm2wat, and wabt-unittests, with warnings treated as errors.
  • All 138 local unit tests pass, including a new seek/backpatch test at offset 4,294,967,296, checking a sparse file size of 4,294,967,297 bytes without allocating a multi-GB buffer. A diagnostic run that deliberately narrowed the seek offset to 32 bits failed this regression as expected; the correct implementation passes.
  • The whole-file Windows read regression stays at 2,147,483,649 bytes because ReadFile intentionally allocates a whole-file buffer; it therefore needs roughly 2GB of memory. It now uses a narrow filename in the current test directory and stdio for file IO, with Windows APIs used only to enable sparse storage. Exclusive file creation prevents overwriting an existing file.
  • Windows x64 CI on cbe0a753 passed, including 137 unit tests and the full test suite. The new regression created a sparse file of 2,147,483,649 bytes, patched at offset 2,147,483,648 through FileStream, then loaded it through ReadFile and verified its size and contents. It ran successfully (not skipped), exercising 64-bit file metadata, seek, tell, and a single application-level fread above 2GB. It requires about 2GB of memory and is restricted to MSVC x64.
  • Formatting and whitespace checks.

Earlier Windows x64 business-sample validation was performed on commit 180b6bc7: a 2,978,116,509-byte WAT file produced a 92,039,608-byte wasm file with exit code 0 and an 8MB stack reserve. That result predates this simplified revision; it is not a new rerun of the business sample.

This fixes large-file IO on Windows x64. It does not promise that a 32-bit process can hold arbitrary multi-GB inputs in memory, or that the parser's whole-file memory requirements have changed.

The MSVC CRT _stat64, _fseeki64, and _ftelli64 interfaces are also available to 32-bit builds; their 64-bit file sizes/offsets are independent of pointer size. The size_t bound remains, so a 32-bit ReadFile buffer cannot represent an input above 4GB. The multi-GB read test is limited to MSVC x64; this revision does not claim a new Windows x86 execution result.

@sbc100

sbc100 commented May 28, 2026

Copy link
Copy Markdown
Member

For the 64-bit file ops can we do something simpler like?

#ifdef _MSC_VER
#define ftell _ftelli64
#define fseek _fseeki64
#endif

For the stack size do you know why we would need such a high value? IIUC the 8mb limit on linux seems to be enough there.

@aydgelee

aydgelee commented May 28, 2026 •

Copy link
Copy Markdown
Author

For the file APIs, I agree the goal is simply to use the 64-bit MSVC CRT variants. I avoided globally redefining ftell / fseek because that changes all uses in the translation unit through macros. Keeping the _fseeki64 / _ftelli64 usage local to the large-file read path makes the change more explicit and avoids affecting unrelated code.

For the stack size, one clarification: /STACK on Windows controls the PE stack reserve size, not memory committed at process startup. It is an upper bound for the thread stack address range, and pages are committed on demand as the stack grows.

The reason this is included here is that, after the 64-bit file I/O fix, the large WAT test no longer failed with value too large; instead, it hit:

0xC00000FD
STATUS_STACK_OVERFLOW

on Windows.

MSVC-linked executables typically default to a 1MB stack reserve, while Linux/macOS environments commonly have larger or configurable stack limits, for example an 8MB soft stack limit on many Linux systems. That likely explains why the same workload can pass on Linux/macOS but fail on Windows.

So the larger /STACK value is intended to raise the Windows stack reserve ceiling for very large WAT parsing workloads. It does not mean the process commits that amount of memory at startup.

@aydgelee
aydgelee force-pushed the fix/windows-large-file-io branch 2 times, most recently from 177c0d9 to e5e08dc Compare May 28, 2026 08:17
@aydgelee

Copy link
Copy Markdown
Author

CI only reported a clang-format issue. No behavior change. Applied the suggested formatting diff.

Comment thread CMakeLists.txt Outdated
@sbc100

sbc100 commented May 28, 2026 •

Copy link
Copy Markdown
Member

For the file APIs, I agree the goal is simply to use the 64-bit MSVC CRT variants. I avoided globally redefining ftell / fseek because that changes all uses in the translation unit through macros. Keeping the _fseeki64 / _ftelli64 usage local to the large-file read path makes the change more explicit and avoids affecting unrelated code.

I'd much rather do it by via redefining these macros, keeping the the code otherwise the same on all platforms.

There really aren't very many file I/O codepaths in wabt, and I imagine we want to be able to handle large files in all of them.

Use MSVC macro redirection for stat, fseek, and ftell so the ReadFile path stays shared across platforms.

Move the MSVC stack reserve setting into the common wabt_executable helper and reduce it to 8MB.
@aydgelee
aydgelee force-pushed the fix/windows-large-file-io branch from e5e08dc to 180b6bc Compare May 29, 2026 10:06
@aydgelee

Copy link
Copy Markdown
Author

Updated.

  • Switched MSVC large-file file I/O to macro-based redirection:
    • stat -> _stat64
    • fseek -> _fseeki64
    • ftell -> _ftelli64
  • Kept the main ReadFile path shared across platforms.
  • Kept the MSVC-only size_t bounds check and chunked fread for 2GB+ inputs.
  • Moved the MSVC stack reserve into wabt_executable().
  • Changed the stack reserve to /STACK:8388608.
  • Non-MSVC behavior is unchanged.

Windows x64 verification:

  • Commit: 180b6bc
  • wasm2wat.exe --version: 1.0.41
  • wat2wasm.exe --version: 1.0.41
  • unity_temp.wat: 2,978,116,509 bytes
  • roundtrip.wasm: 92,039,608 bytes
  • wat2wasm exit code: 0

No "value too large", no "unable to read file", and no STATUS_STACK_OVERFLOW with the 8MB stack reserve.

Comment thread src/stream.cc Outdated
@sbc100

sbc100 commented May 29, 2026

Copy link
Copy Markdown
Member

Why do we need to read and write the files in chunks on windows? But not on other platforms? Could you update the PR description to include that maybe?

@sbc100

sbc100 commented May 29, 2026

Copy link
Copy Markdown
Member

Are you seeing these error on 64-bit windows too? Or is this mostly a fix for 32-bit windows?

Nishuuzz added a commit to Nishuuzz/wabt that referenced this pull request Aug 9, 2026
CI showed 512 still faults on windows: msvc's default stack is 1MB,
which is not enough to reach the limit, so deeply nested input dies
before the parser can report anything.

Picking a limit that does fit in 1MB is not really an option. It would
have to be somewhere under 256, and real modules need more than that --
round-tripping sqlite needs 320 -- so it would start rejecting input
that parses fine today on Linux and macOS.

Asking for the 8MB those platforms already give us instead. This is the
same change WebAssembly#2748 makes for the same reason; happy to drop it here if
that lands first.
Use MSVC 64-bit file APIs with shared stdio paths. Remove the redundant SeekFile and chunked IO helpers after checking the UCRT implementation. Guard non-MSVC seek offsets and add Windows x64 coverage above 2GB. Keep the upstream 8MB stack setting.
@aydgelee aydgelee changed the title Fix Windows/MSVC handling of large WAT input files and increase stack reserve Fix Windows/MSVC handling of large files Oct 2, 2026
@aydgelee

aydgelee commented Oct 2, 2026 •

Copy link
Copy Markdown
Author

Updated the code and rewrote the PR description to match it.

On the chunking question: I checked Microsoft's published UCRT source package. stdio/fread.cpp already bounds its low-level reads to INT_MAX, and stdio/fwrite.cpp also splits large requests. The _read limit therefore does not justify an extra limit on a whole fread call. I removed both additional Windows-only chunked IO paths; the revision now keeps the normal fread / fwrite calls shared across platforms.

On the platform question: the reported failure and the earlier successful 2,978,116,509-byte WAT conversion were both on Windows x64, not a 32-bit Windows process. MSVC's long and its default file metadata/seek/tell interfaces still have 32-bit limits there; _stat64, _fseeki64, and _ftelli64 address that separately from process pointer size. The size_t check remains for buffers that cannot represent the file size.

I also merged current main, resolved the ByteSpan changes, and retained the upstream 8MB stack setting from #2815. The old 512MB description has been removed.

Local warnings-as-errors builds and all 137 unit tests pass. All 17 GitHub checks on the final commit cbe0a753 have now passed. Windows x64 CI passed the build, 137 unit tests, C API tests, and full test suite. The new regression successfully read a 2,147,483,649-byte sparse file and patched it beyond the 2GB boundary, specifically checking the simplified, unchunked IO path; it ran rather than being skipped. The test is explicitly gated on MSVC x64 to match the scope of the fix. The earlier business-sample result remains historical; I am not claiming it was rerun for this revision.

@aydgelee

aydgelee commented Oct 2, 2026 •

Copy link
Copy Markdown
Author

@sbc100 The revisions are ready for another review at cbe0a753.

The earlier stack-size and SeekFile threads have been addressed and marked resolved. The PR now uses macro-based MSVC file API redirection and shared stdio paths, with the redundant Windows-only chunking removed. The Windows x64 scope and UCRT findings are documented in the updated description.

All 17 GitHub checks passed, including the Windows x64 regression that reads and patches a file above 2GB. Could you take another look when convenient?

@sbc100 sbc100 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm with couple of question about testing.

@zherczeg WDYT?

Comment thread src/test-stream.cc Outdated
Comment thread src/test-stream.cc Outdated

@zherczeg zherczeg left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It is a bit of unclear the target of this patch for me. Do we need to process > 4GB files? Is this a real use case?

Comment thread src/common.cc Outdated
Comment thread src/common.cc
@aydgelee

aydgelee commented Oct 5, 2026 •

Copy link
Copy Markdown
Author

@zherczeg Yes, this is a real use case: I have used input files larger than 4GB, and that workload is how I encountered the bug. The previously documented 2,978,116,509-byte sample is an earlier validation example, not the largest input I have needed to process.

The target is large-file IO on Windows/MSVC, including a 64-bit process, where the default file-size and seek/tell interfaces still use 32-bit sizes or offsets. This patch is not a streaming-parser redesign and does not promise that a 32-bit process can hold a multi-GB input in memory.

The follow-up is now pushed as 2d6c0725, including current upstream main. The temporary-file test setup and duplicate MSVC include block have been simplified, and the API-availability and memory questions have been answered in their respective threads.

I added a sparse-file seek/backpatch regression above 4GB that does not allocate a 4GB buffer. The existing whole-file read regression stays just above the signed 32-bit boundary to limit its memory cost. The larger real-world input is a reported use case, not a claim that I have rerun that input in CI.

Windows x64 CI passed the build, all 138 unit tests, C API tests, and full test suite. Both the >4GB seek/backpatch regression and the revised 2GB+ read regression ran successfully, rather than being skipped. Other checks are still completing. @sbc100 The testing follow-up is ready for another look.

@zherczeg zherczeg left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Comment thread src/test-stream.cc

#if !COMPILER_IS_MSVC
TEST(FileStream, RejectUnrepresentableSeekOffset) {
std::unique_ptr<FILE, FileCloser> file(tmpfile());

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I guess maybe on linux/UNIX we should use fseeko and set _FILE_OFFSET_BITS=64? But that can be followup of course

@sbc100

sbc100 commented Oct 5, 2026

Copy link
Copy Markdown
Member

Can you suggest a short summary for the commit message for this change, either by updating the PR description of posting your proposed commit message here as a comment.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants