ML researcher · Mountain View, California

Yash
Maheshwari.

I build small language models that perform better with fewer resources, write the benchmark tasks that frontier agents still fail, and bring a student's view of AI to classrooms, conference stages and state legislatures.

Portrait of Yash Maheshwari
NowHealthcare AI · Stanford Shah Lab
  • EMNLP ’26First-author paper, accepted to the BabyLM Workshop
  • 2 Stanford labsShah Lab (Medicine) and Lemons Lab (Education)
  • 2 patentsAI agent infrastructure filings, 2025
  • 9,000+Expected audience for my FETC 2027 keynote

Research

01 / 05
Accepted · EMNLP 2026 BabyLM WorkshopFirst author

Halved CLM exposure: keeping small recurrent models from losing grammar late in training

Question
Why do small recurrent language models get worse at grammar late in training?
Model
RWKV-7, a 27.4M-parameter O(T) recurrent model, trained locally on a MacBook in MLX
Result
Halving how often the model is updated curbs a 2.6 to 3.9 point late drop in grammar score
Benchmark
Beats the official GPT-2 baseline on BLiMP with about 28% of its parameters
Grammar score (BLiMP), official BabyLM 2026 pipeline
Our model · 27.4M68.70%
GPT-2 baseline · 98.4M65.23%

Bar length = parameter count

Accepted · NeurIPS 2026 ODI Workshop (poster)AXIOM paper under review

Below one bit: how you train a tiny model matters more than its size

  • AXIOM (under review) · Training Into the Container: The Projection GapCompressing a trained model below 1 bit per weight breaks it. Retraining just 8% of it recovers 97%.
  • ODI (accepted) · Training Route and Budget Shape RobustnessNatively trained sub-1-bit models handle noisy input better, on every seed and scale tested.
Prediction error at 0.41 bits/weight · lower is better
Trained natively~3.1 bpb
Compressed after training5.8–6.4 bpb

Past 5.06 bpb, a model is worse than guessing letter frequencies

  • In developmentHealthAdminBench V2 · hospital-admin tasks frontier agents still fail, on a realistic Epic EHRStanford Medicine × Kinetic
  • AcceptedSmall models as math tutors · fine-tuned 8B Llama hits 93.7% step accuracyJEI · first author · in press
  • Patent · filedHierarchical aggregation tree for MCP server selection · routes agent requests to the right toolNon-provisional · sole inventor · Oct 2025
  • Patent · filedPredictive compliance for AI agents · flags unsafe tool calls before they runProvisional · co-inventor with Aisera · Oct 2025

All accepted publications, with abstracts and PDFs →

Experience

02 / 05
  1. Aug 2026 — PresentStanford Medicine, Shah Lab

    Research Intern, Healthcare AI

    • Maintaining HealthAdminBench; A/B testing models with AI lab engineers
    • Building HealthAdminBench V2 on Harbor + a realistic Epic EHR
    • Tracing where each model gets stuck · with Dr. Nigam Shah
    Epicreal EHR environment
  2. Jul — Aug 2026Kinetic Systems

    AI Research Contractor

    • Wrote healthcare workflow tasks frontier models could not yet solve
    • Traced model error paths on each task for frontier-lab clients
    FrontierAI lab clients
  3. Jun 2025 — PresentStanford Graduate School of Education, Lemons Lab

    Research Intern: Kai & PAWS

    • Kai: AI reading tutor in 10+ districts, 50+ teachers
    • Built the AIOps eval suite for Kai · cut latency 75%, then a non-LLM path answering in <50 ms
    • PAWS: kindergarten handwriting tutor, web prototype → native iPad app
    1,200+students on Kai
  4. Jun — Aug 2025Aisera

    AI Engineering Intern

    • MCP servers linking agents to Salesforce, Clari and Slack
    • Open-source MCP bridge for HTTP + SSE clients
    • Co-inventor on 1 provisional patent · 2nd in company hackathon
    4MCP servers developed
  5. Sep 2024 — PresentMountain View High School

    Executive Board Member, Principal's Tech Internship

    • Joined at launch · Executive Board since Jun 2025
    • Shipped AI Policy Pathway and Bridge the Gap 360
    • Organized Parent Night, CAL-MSCS Day and an AI Playlab
    2tools shipped

On stage

03 / 05
“Misaligned incentives are dangerous.”
From my Common Sense Media 2026 talk
Featured & interviewed by

The Washington Post · Los Altos Town Crier · Center for Digital Education · Amplify · Thinkering Collective

  1. 2026
    Common Sense Media Summit · Two keynote panelsOpened and closed day one · met Secretary Hillary Clinton
    600+
  2. 2026
    FETC · Two sessionsAI ethics through play · student-led tech internships
    Orlando
  3. 2025
    ASU+GSV Summit · “Learners Light the Way”AI Show demo · 1 of 3 high schoolers at Walton breakfast
    San Diego
  4. 2025
    Google HQ · Panel with engineers and designersHow students use AI, and what to design for
    Mountain View
  5. Also: Foothill College KCI · CAL-MSCS statewide educator day (80+ teachers) · AI & Education Parent Night (60+) · AI Playlab (100+) · Stanford Down Syndrome Conference

AI policy

04 / 05
  • 1,500+legislators contacted
  • 75+AI bills tracked
  • 10legislative offices
AI bills discussed with legislative offices, 2026
StateBillTopic
FLSB 482AI Bill of Rights
MELD 2162Human-like features in AI
MISB 760AI companion chatbot regulation
PASB 939 · SB 1090AI regulatory sandbox, AI safeguards
COHB26-1139AI in health care
NYA09253AI use in policing
OHSCR 14State authority over AI regulation
MOHB 2239Data center buildout
CA—Student perspective, office of Rep. Sam Liccardo

Founded & awarded

05 / 05

Founded & led

  • RL Game ClubFounder · RL through game-bot competitions
  • Tech Spark 501(c)(3)Co-founder · 5 summers of K-8 robotics and coding
  • FTC robotics teamCo-founder & student mentor
  • FRC 9584Software lead · 27th in FRC Championship Division

Awards

  • MVHacks, 1st place overall2025
  • Aisera AI Hackathon, 2nd of 30 teams2025
  • Highest Rookie Team Award, FRC Worlds2024
  • Congressional App Challenge, Honorable Mention2024
  • FTC Judges' Choice Award2024–25