# Yash Maheshwari > ML researcher (high school, Mountain View, California, Class of 2027) working on efficient small language models, healthcare AI agent benchmarks, and student-led AI policy. Research intern at Stanford's Shah Lab (Stanford Medicine) and Lemons Lab (Stanford Graduate School of Education). ## Research - **Halved CLM Exposure Mitigates Late-Training Degradation in Small Recurrent Language Models** (accepted, archival, BabyLM Workshop at EMNLP 2026; sole author): halving how often causal language modeling updates the weights mitigates late-training BLiMP degradation in small RWKV-7 (27.4M-parameter) recurrent models trained locally with MLX. Official BabyLM 2026 pipeline: 68.70% BLiMP, above the GPT-2 baseline (65.23%, 98.4M parameters) at about 28% as many parameters. Page: https://research.yash-maheshwari.com/papers/babylm-2026 - **Training Into the Container: The Projection Gap Below One Bit per Block-Linear Weight** (submitted, NeurIPS 2026 AXIOM Workshop): below one bit per weight, one-shot post-training compression lands at 5.8-6.4 bits per byte while native quantization-aware training stays near 3; adapting only the full-precision interface (~8% of parameters) recovers 97% of the gap. - **Below One Bit: Training Route and Budget Shape Robustness Under On-Device Storage Constraints** (accepted, poster, NeurIPS 2026 ODI Workshop; sole author): natively trained sub-1-bit models degrade less under corrupted input on code, at every scale and seed tested, but the contrast depends strongly on the token budget. Page: https://research.yash-maheshwari.com/papers/odi-2026 - **HealthAdminBench V2** (in development, Stanford Shah Lab with Kinetic Systems): benchmark of LLM agents on healthcare administrative work (prior authorizations, denial appeals, DME orders) in a realistic Epic EHR environment on the Harbor framework. - **Training and comparison of fine-tuned small language models as math tutoring assistants** (accepted, in press, Journal of Emerging Investigators; first author, with Rinky Gupta): fine-tuned Llama tutors at 1B, 3B and 8B parameters; parameter count mattered more than training epochs; the 8B model reached 93.65% step accuracy. Page: https://research.yash-maheshwari.com/papers/jei-2026 - **Patents (filed 2025)**: Hierarchical Aggregation Tree for MCP Server Selection and Execution (non-provisional, sole inventor); Predictive Compliance for AI Agents (provisional, co-inventor with Aisera). ## Experience - Research Intern, Shah Lab, Stanford Medicine (Prof. Nigam Shah), Aug 2026-present; AI Research Contractor, Kinetic Systems, Jul-Aug 2026. - Research Intern, Lemons Lab, Stanford GSE (Prof. Chris Lemons), Jun 2025-present: Kai AI reading tutor (1,200+ students, 50+ teachers, 10+ districts); PAWS handwriting iPad app. - AI Engineering Intern, Aisera, Jun-Aug 2025: MCP servers for Salesforce, Clari and Slack; MCP bridge for HTTP and SSE. - Executive Board Member, MVHS Principal's Tech Internship, Sep 2024-present. ## Speaking and policy - Keynote, FETC 2027 (Orlando, 9,000+ expected). Two keynote panels, Common Sense Media Summit 2026. Two sessions at FETC 2026. ASU+GSV Summit 2025. Panel at Google HQ. - Featured in The Washington Post (Oct 2025) for co-designing a school district AI philosophy with students, parents and teachers. - Met 10 legislative offices across 9 states on AI bills; contacted 1,500+ state legislators on 75+ bills. ## Links - Website: https://www.yash-maheshwari.com/ - Publications (accepted work, abstracts, PDFs): https://research.yash-maheshwari.com/ - Resume (PDF): https://www.yash-maheshwari.com/resume/Yash_Maheshwari_Resume.pdf - GitHub: https://github.com/yashmahe2020 - LinkedIn: https://www.linkedin.com/in/yashmaheshwari2009/ - Email: yashmahe2018@gmail.com