Information Analysis · University of Michigan '27

I check what a number is actually counting before I trust it.

I build AI systems and ship them in public: two so far, each with a write-up and measured results. Every result in my AI projects came from a measured run.

Judd Gurtman

AI projects

The Year of AI: one project at a time, shipped in public, each with a write-up of what broke.

Research agent

Python · Claude API · web search · citations

Ask it a hard question and it breaks the question into smaller ones, researches them in parallel, checks every claim against the exact quote it came from, and writes a short brief where every claim has to cite its source.

  • 64% → 41%weakly supported claims, after measuring why they were weak before fixing anything
  • 45 / 46sentences cited, 0 invented sources
  • $1.68 → $1.25per question, from a per-step cost ledger

Agent fleet

MCP · orchestration

One orchestrator agent that runs my other agents as tools, through an MCP server. Anything that spends money has a cap and waits for a person to approve it, and a stuck job gets shut down so it stops costing money.

Coach

FastAPI · tool use · PWA

An AI personal trainer web app that takes real actions: it logs workouts, meals, weigh-ins and plans through 7 tools and builds each week from what you actually did. Two QA passes found and fixed 25 bugs.

Data projects

Fuel price volatility

Python · pandas · EIA data

30+ years of weekly U.S. gasoline and diesel prices (26,000+ observations): how volatile they are, how regional prices drift apart, and which shocks drove the biggest swings.

Movie ratings

Python · APIs · SQLite

A pipeline that pulls film data from the TMDB and OMDb APIs into SQLite, then looks at how ratings and box office vary by genre.

Experience

Data Analytics Intern

TrueSource, an OnPoint Group company · June to July 2026 · remote

  • My first estimate of a system's failure rate looked alarming. I took it to the person who owns the process, learned that much of what I'd counted was routing working as designed, and re-cut the outcomes into intended routing, caller error and true system error. The real error rate was far lower and matched the team's own figure.
  • Found why the data was confusing: many call records weren't linked to a Salesforce work order, so outcomes couldn't be measured. Proposed a first-contact resolution metric, and learned it needs a clear definition before anyone builds scorecards on it.
  • Audited 29 KPIs across 4 operational dashboards and documented 7 inconsistencies between them, from formulas and date filters. The analytics team added disclaimers to the live dashboards.
  • Validated a Power BI model against Salesforce. After my manager's review I changed my first verdict and documented what my first check missed.
  • Used AI to classify a large set of calls by reason and resolution, and added a third angle to the analysis that nobody had asked for.
  • Gave five presentations in six weeks, the last one to company leadership.

This is where "check what a number is actually counting" comes from.

For fun

Things I build because I want them to exist.

Clout Royale home base: a pixel-art mansion with upgrade stations Clout Royale mission: top-down arena fight against parody creator characters

Clout Royale

TypeScript · Phaser 3

A satirical top-down arena game about the creator economy. You play a gym influencer defending your mansion from waves of parody creators. Eliminations earn clout, your Aura meter powers your abilities, and your Cringe meter slows you down when you overdo it.

Desktop only: keyboard and mouse.

How I work

  1. Measure before fixing.

    My research agent's obvious fix was aimed at the wrong cause. Counting first showed the real one.

  2. Only claim what I measured.

    Every AI result here traces back to a run, a trace file or a controlled replay.

  3. A person approves anything risky.

    If an agent is about to spend money or talk to someone, it waits for a human.

About

I'm a senior at the University of Michigan studying Information Analysis, graduating in May 2027. I like the part of analytics where the data is messier than the dashboard suggests, and I've spent this year learning to build AI that holds up there.

Outside of class I coach youth flag football, which I've done as a volunteer since January 2024. I grew up in Aspen playing travel lacrosse, and this spring I was an assistant varsity lacrosse coach. I ski whenever I can, follow sports analytics a little too closely, and I'm into house music, which I want to learn to produce.