Skip to content
View Tusharv23's full-sized avatar
  • Wood Mackenzie

Block or report Tusharv23

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Tusharv23/README.md

Tushar Verma

Professionally suspicious of what AI agents claim to remember.

LLM agent-memory & reliability evals  ·  I break memory systems and file the receipts.

Python LLM Evals Agent Memory


What I work on

I evaluate the memory layer of AI agents — the part that decides which facts to keep, replace, or protect over time. Most eval work checks whether a model retrieves the right thing. I check whether it changes its mind for the wrong reasons: overwriting a confirmed fact on weak evidence, softening a safety-critical one, or confabulating a merge of two conflicting memories.

My approach is deliberately unglamorous — small, adjudicated, honest samples with error bars, and claims that get killed when the intervals overlap.

Selected work

  • conflict-resolution-evals — a cross-model evaluation of a production agent-memory policy. 120 manually adjudicated runs across 10 models; a follow-up context-load experiment (216 runs) with Wilson confidence intervals. Found action-type memory edits robust but restraint-type edits (don't-overwrite-on-weak-evidence) failing ~50% of the time — including a confirmed severe allergy silently downgraded. Findings were reported upstream to the open-source framework.

How I think about this

  • Negative results are results. Several findings in the study above did not survive uncertainty quantification, so I retired them.
  • A green check is only as good as its assertion.
  • If I can't show the arithmetic, I don't trust the number — including my own.

Elsewhere

  • LinkedIn: tushar-verma  ·  writing about evals & agent memory
  • Background: platform engineering (API gateways, multi-team infra at scale), now focused on AI reliability.

Pinned Loading

  1. conflict-resolution-evals conflict-resolution-evals Public

    Python

  2. heauristic-sampler heauristic-sampler Public

    Python