Events
Free live lessons on LLM evaluation and data science. Each session is practical, code-first, and designed to leave you with something you can use the same day.
Upcoming Events
Sat Oct 16 · 02:00 PM–6:00 PM EDT · $300
Build a Production-Grade LLM Eval Harness
In four hours, build a production eval harness from scratch — calibrated judges, error bars on every metric, paired tests that settle arguments, and a CI gate that blocks bad merges. You leave with a running Inspect-AI harness, the full repo, and a swap guide to run it on your own product data. Monday morning, it runs.
Discount code
MEASURE20
Maven Workshop · Live on Zoom · 4 hours
Enroll — $300 →Wed Oct 21 · 12:00 PM ET · Free · 30 min
Jev in One Call
See Jev's three primitives on a live support ticket: Choice picks the team, Score sets the urgency, Noul answers the refund. Fan out a dozen questions in one round trip and read the 100ms latency and cost lines that come back with every call.
Maven · Free lightning lesson · 30 min · Zoom + recording
Register free →
Sat Oct 31 · 9:30 AM EDT · $149
Harness Engineering: Build an AI Data Analyst You Can Actually Trust
Join Bruno Goncalves for a 4-hour, hands-on virtual workshop to engineer a reliable, production-ready AI data analyst in Python. You will connect an agent to permissioned MCP tools, manage context within a hard token budget, build layered guardrails against prompt injection, verify the numbers in an AI report, and make failures diagnosable with retries, budgets and structured traces — ending with a reusable, framework-free Python agent harness. Runs on your laptop: no GPU, cloud account, or API key required.
Packt Publishing · Live on Luma · 4 hrs
Enroll — $149 →Wed Nov 4 · 12:00 PM ET · Free · 30 min
Can You Trust Jev's scores?
Pull 113 labeled yes/no answers, draw the reliability diagram for Jev and for Claude's self-reported confidence, compute ECE and Brier score in ten lines of numpy, and set auto / confirm / human bands from the cost of a false allow vs a false block.
Maven · Free lightning lesson · 30 min · Zoom + recording
Register free →Wed Nov 11 · 12:00 PM ET · Free · 30 min
Gate Your Agent in 100 ms
Gate 41 proposed agent actions (five of them adversarial) with Jev and Claude against the same written policy. Measure recall on unsafe actions, false-block rate, p95 latency, and cost per million checks — then set the auto / confirm / human bands.
Maven · Free lightning lesson · 30 min · Zoom + recording
Register free →Wed Nov 18 · Live cohort · 4 hours
Jev vs Claude: A Measured Guide
Learn Jev end to end, then prove task by task where it beats Claude or not. Build a paired bakeoff harness, run it against Claude Opus 5.5 on routing, scoring, judging and agent gating, and leave with the confidence intervals, the cost per million decisions, and an eleven-question decision guide.
Discount code
JEV2020
Maven · Live cohort · 4 hours
Enroll — $300 →Past Events
Wed Oct 7 · 2:00 PM EDT · Free
Stop Bad Merges with an LLM Eval Gate
A bad prompt change reaches production the same way a good one does: nobody measured either. In this free lesson, learn to wire an eval suite into GitHub Actions so regressions die in the pull-request queue — not in front of users. The full setup fits in one YAML file. Vigilance does not scale. Policy does.
Maven · Live on Zoom · 30 min
Wed Sep 23 · 2:00 PM EDT · Free
Put Error Bars on Your LLM Metrics
You ran the same eval twice and got 84.2, then 81.9. Which number goes in the report? Without error bars, every score is a coin flip dressed as a fact. In 30 minutes, learn to bootstrap a confidence interval on any metric — accuracy, cost, or latency — in 20 lines of Python. No distribution assumptions required.
Maven · Live on Zoom · 30 min
Sat Sep 12 · 9:30 AM–1:00 PM EDT
Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps
A hands-on 3.5-hour deep dive into production LLM engineering — covering evals, RAG, agents, prompt engineering, observability, and LLMOps. Every technique demonstrated with running code. Organized by Packt Publishing.
Packt Publishing · Live online · 3.5 hours
Wed Sep 9 · 2:00 PM EDT · Free
Prove Your Prompt Change Actually Helped
You changed a prompt. The score went up. But did it actually improve — or did you get lucky? In this free 30-minute lesson, learn how to run a proper paired test on two prompt versions, how to tell a real improvement from noise, and how to pick a sample size that settles the argument. No more gut-feel comparisons.
Maven · Live on Zoom · 30 min
Fri Jul 11
Production Graph RAG: Build Explainable LLM Apps with Knowledge Graphs
A deep dive into Graph RAG — combining knowledge graphs with retrieval-augmented generation to build LLM applications that are explainable, auditable, and grounded in structured data. Organized by Packt Publishing.
Packt Publishing · Live online
Tue Jul 8 · 1:00–5:00 PM EDT
Automate the Boring Developer Stuff with LLMs
A four-hour hands-on workshop covering how to use LLMs to automate repetitive developer tasks — code generation, test writing, documentation, and more. Every exercise ships with working code.
O'Reilly Live Training · 4 hours