Knowledge & memory
What an agent actually knows, not just what it retrieved once. Retrieval is a lookup; knowing is what survives the next turn.
NIKHIL KADAPALA / PHD STUDENT · UNH
I build AI systems that have to work on messy, real-world text, then I try to measure whether they actually do. My PhD is on agent evals: knowledge bases, retrieval, and memory.
01 / THE ARGUMENT
A high score can describe a system nobody can use. I keep running into the same gap: the metric goes up, and the thing a person actually needed still isn’t there. That distance — between a plausible answer and a useful one — is most of what I work on.
How I approach itPUBLISHED RESEARCH
We compared fine-tuning and prompting setups for claim extraction. FLAN-T5 won on METEOR, while iterative self-refinement sometimes produced claims a fact-checker would actually want to work with. That gap is the whole plot: evaluation has to stay connected to the user and the work the system is meant to support.
SELECTED SYSTEMS
A career copilot for fit, preparedness, and alignment—not another apply-to-everything button.
Sandbox and harness primitives for coding agents. Zero dependencies, on purpose.
Multimodal agentic RAG with a real evaluation harness. “It felt pretty good” is not a metric.
Noisy social posts to concise, checkable claims—class project to paper to pip install.
CURRENTLY INTO
What an agent actually knows, not just what it retrieved once. Retrieval is a lookup; knowing is what survives the next turn.
Score the behavior against intent, not a storyboarded workflow. A golden trace is a useful fixture and a terrible definition of success.
Disaggregated prefill/decode, speculative decoding, prefix caching. Cheaper and faster serving, without pretending that's the same as a smarter model.
SAY HELLO
Always happy to talk about research, agents, or evaluation that refuses to behave. LinkedIn or X is the easiest ping.