-
Notifications
You must be signed in to change notification settings - Fork 350
Pull requests: microsoft/onnxruntime-genai
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Support DeepStack (Qwen3-VL) in the vision->embedding->decoder pipeline
#2513
opened Sep 1, 2026 by
Tachion (SanjayAMD)
Contributor
Loading…
Validate Qwen vision embedding output
#2511
opened Aug 31, 2026 by
Akshay Sonawane (apsonawane)
Contributor
Loading…
[WebGPU] fixes pointer arithmetic bug for multi-image inputs for mistral3
#2510
opened Aug 31, 2026 by
Prathik Rao (prathikr)
Contributor
Loading…
Add model-independent decoded stop-string matching
#2509
opened Aug 31, 2026 by
Baiju Meswani (baijumeswani)
Collaborator
Loading…
Harden Engine request lifecycle failures
#2507
opened Aug 31, 2026 by
Baiju Meswani (baijumeswani)
Collaborator
Loading…
Set top_k to vocab_size in top-p tests
#2505
opened Aug 31, 2026 by
Hrishith Thadicherla (hthadicherla)
Contributor
Loading…
feat: accept oci:// model sources in the model builder
#2503
opened Aug 30, 2026 by
Eric Curtin (ericcurtin)
Loading…
1 task done
Reuse the block drafter runtime for DSpark
#2497
opened Aug 28, 2026 by
Tianlei Wu (tianleiwu)
Contributor
•
7/7
Loading…
4 of 5 tasks
Run a DFlash2 block drafter in the batch engine
#2496
opened Aug 28, 2026 by
Tianlei Wu (tianleiwu)
Contributor
•
6/7
Loading…
5 tasks done
Run chained MTP drafting in the batch engine
#2495
opened Aug 28, 2026 by
Tianlei Wu (tianleiwu)
Contributor
•
5/7
Loading…
4 tasks done
Add support for quant_auto format in model builder
#2492
opened Aug 28, 2026 by
Vishal Jain (VishalX)
Contributor
Loading…
Export DSpark block drafters for Qwen 3.5 and 3.8
#2490
opened Aug 28, 2026 by
Tianlei Wu (tianleiwu)
Contributor
Loading…
Generalize Gemma 4 MoE Quark path to plain Quark/AWQ int4 experts
#2483
opened Aug 27, 2026 by
Thiago Pereira Rocha (thpereir)
Contributor
Loading…
3 tasks done
Add 2-bit (uint2) Gemma 4 MoE support to the model builder
#2482
opened Aug 27, 2026 by
Thiago Pereira Rocha (thpereir)
Contributor
Loading…
3 tasks done
[C#] Expose cached Generator through IChatClient.GetService
#2475
opened Aug 27, 2026 by
ZedingZhang (ZedingZhang)
Loading…
Add Gemma 4 text model builder support (dense gemma-4-12b-it + 26B-A4B MoE)
#2473
opened Aug 26, 2026 by
Thiago Pereira Rocha (thpereir)
Contributor
Loading…
Support stateful KV cache for multimodal decoder
#2469
opened Aug 25, 2026 by
Anirudh Swaminathan (Anirudh-Swaminathan)
•
Draft
[benchmark] Add prompt token truncation and trim benchmark output
#2468
opened Aug 25, 2026 by
Namal Rajatheva (amd-namalr)
Loading…
Harden Engine request lifetime and sampler acquisition
#2448
opened Aug 22, 2026 by
bmehta001
Contributor
Loading…
Keep Python opaque data alive for the lifetime of the request
#2445
opened Aug 21, 2026 by
Gopalakrishnan Nallasamy (GopalakrishnanN)
Contributor
Loading…
Warn when building gpt-oss with float16 I/O precision
#2444
opened Aug 21, 2026 by
Gopalakrishnan Nallasamy (GopalakrishnanN)
Contributor
Loading…
Fix AttributeError in KV cache shape naming under paged attention
#2443
opened Aug 21, 2026 by
Gopalakrishnan Nallasamy (GopalakrishnanN)
Contributor
Loading…
Previous Next
ProTip!
Adding no:label will show everything without a label.