Skip to content

[Bug] Vision API VRAM Leak (onnxruntime-genai) #1034

Description

@A795da

Describe the issue

Vision API VRAM Leak (onnxruntime-genai)

When running vision queries through the qwen3.5-4b model, the underlying engine (onnxruntime-genai) does not fully release VRAM between inference calls ; it climbs to ~11.9 GB over 3 requests, latency escalates from 54s→90s, then crashes with a CUDA error in CudaMallocArray...out of memory.

-Impact: Runtimes escalated severely (from 54 seconds to 90 seconds) before crashing entirely with a CUDA error in CudaMallocArray... out of memory.

To reproduce

Run 3+ consecutive vision queries against the vision model without unloading(qwen3.5-4b). Compare against the workaround queue in local-rag/src/vision-model-guard.js, which force-unloads/reloads every 2 requests.

Urgency

High causes a hard crash in productio.

Platform and architecture

Windows X64

OS Version

Windows 11

Installation type

Released package/binary

Foundry Local version

0.10.2

API or surface area

Other / Unknown

Hardware acceleration/backend

CPU

Backend/runtime version

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions