Events

Free live lessons on LLM evaluation and data science. Each session is practical, code-first, and designed to leave you with something you can use the same day.

Upcoming Events

Build a Production-Grade LLM Eval Harness
$300

Sat Oct 16 · 02:00 PM–6:00 PM EDT · $300

Build a Production-Grade LLM Eval Harness

In four hours, build a production eval harness from scratch — calibrated judges, error bars on every metric, paired tests that settle arguments, and a CI gate that blocks bad merges. You leave with a running Inspect-AI harness, the full repo, and a swap guide to run it on your own product data. Monday morning, it runs.

Discount code

MEASURE20

Maven Workshop · Live on Zoom · 4 hours

Enroll — $300 →
Jev in One Call
Free

Wed Oct 21 · 12:00 PM ET · Free · 30 min

Jev in One Call

See Jev's three primitives on a live support ticket: Choice picks the team, Score sets the urgency, Noul answers the refund. Fan out a dozen questions in one round trip and read the 100ms latency and cost lines that come back with every call.

Maven · Free lightning lesson · 30 min · Zoom + recording

Register free →
Harness Engineering: Build an AI Data Analyst You Can Actually Trust
$149

Sat Oct 31 · 9:30 AM EDT · $149

Harness Engineering: Build an AI Data Analyst You Can Actually Trust

Join Bruno Goncalves for a 4-hour, hands-on virtual workshop to engineer a reliable, production-ready AI data analyst in Python. You will connect an agent to permissioned MCP tools, manage context within a hard token budget, build layered guardrails against prompt injection, verify the numbers in an AI report, and make failures diagnosable with retries, budgets and structured traces — ending with a reusable, framework-free Python agent harness. Runs on your laptop: no GPU, cloud account, or API key required.

Packt Publishing · Live on Luma · 4 hrs

Enroll — $149 →
Can You Trust Jev's scores?
Free

Wed Nov 4 · 12:00 PM ET · Free · 30 min

Can You Trust Jev's scores?

Pull 113 labeled yes/no answers, draw the reliability diagram for Jev and for Claude's self-reported confidence, compute ECE and Brier score in ten lines of numpy, and set auto / confirm / human bands from the cost of a false allow vs a false block.

Maven · Free lightning lesson · 30 min · Zoom + recording

Register free →
Gate Your Agent in 100 ms
Free

Wed Nov 11 · 12:00 PM ET · Free · 30 min

Gate Your Agent in 100 ms

Gate 41 proposed agent actions (five of them adversarial) with Jev and Claude against the same written policy. Measure recall on unsafe actions, false-block rate, p95 latency, and cost per million checks — then set the auto / confirm / human bands.

Maven · Free lightning lesson · 30 min · Zoom + recording

Register free →
Jev vs Claude: A Measured Guide
$300

Wed Nov 18 · Live cohort · 4 hours

Jev vs Claude: A Measured Guide

Learn Jev end to end, then prove task by task where it beats Claude or not. Build a paired bakeoff harness, run it against Claude Opus 5.5 on routing, scoring, judging and agent gating, and leave with the confidence intervals, the cost per million decisions, and an eleven-question decision guide.

Discount code

JEV2020

Maven · Live cohort · 4 hours

Enroll — $300 →

Past Events

Stop Bad Merges with an LLM Eval Gate
Past Free

Wed Oct 7 · 2:00 PM EDT · Free

Stop Bad Merges with an LLM Eval Gate

A bad prompt change reaches production the same way a good one does: nobody measured either. In this free lesson, learn to wire an eval suite into GitHub Actions so regressions die in the pull-request queue — not in front of users. The full setup fits in one YAML file. Vigilance does not scale. Policy does.

Maven · Live on Zoom · 30 min

Put Error Bars on Your LLM Metrics
Past Free

Wed Sep 23 · 2:00 PM EDT · Free

Put Error Bars on Your LLM Metrics

You ran the same eval twice and got 84.2, then 81.9. Which number goes in the report? Without error bars, every score is a coin flip dressed as a fact. In 30 minutes, learn to bootstrap a confidence interval on any metric — accuracy, cost, or latency — in 20 lines of Python. No distribution assumptions required.

Maven · Live on Zoom · 30 min

Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps
Past $150 speaker

Sat Sep 12 · 9:30 AM–1:00 PM EDT

Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps

A hands-on 3.5-hour deep dive into production LLM engineering — covering evals, RAG, agents, prompt engineering, observability, and LLMOps. Every technique demonstrated with running code. Organized by Packt Publishing.

Packt Publishing · Live online · 3.5 hours

Prove Your Prompt Change Actually Helped
Past Free

Wed Sep 9 · 2:00 PM EDT · Free

Prove Your Prompt Change Actually Helped

You changed a prompt. The score went up. But did it actually improve — or did you get lucky? In this free 30-minute lesson, learn how to run a proper paired test on two prompt versions, how to tell a real improvement from noise, and how to pick a sample size that settles the argument. No more gut-feel comparisons.

Maven · Live on Zoom · 30 min

Production Graph RAG: Build Explainable LLM Apps with Knowledge Graphs
Past Paid

Fri Jul 11

Production Graph RAG: Build Explainable LLM Apps with Knowledge Graphs

A deep dive into Graph RAG — combining knowledge graphs with retrieval-augmented generation to build LLM applications that are explainable, auditable, and grounded in structured data. Organized by Packt Publishing.

Packt Publishing · Live online

Automate the Boring Developer Stuff with LLMs
Past Paid

Tue Jul 8 · 1:00–5:00 PM EDT

Automate the Boring Developer Stuff with LLMs

A four-hour hands-on workshop covering how to use LLMs to automate repetitive developer tasks — code generation, test writing, documentation, and more. Every exercise ships with working code.

O'Reilly Live Training · 4 hours

​

Subscribe to get our latest content by email.
    We won't send you spam. Unsubscribe at any time.