Event Highlights
LLM-as-a-judge vs. human-in-the-loop oversight
Join our CTO and Co-Founder, Hui Zhang, at VB Transform, for the great debate on model validation.
As agents move from prototype to production, the evaluation bottleneck has become one of the biggest hurdles. Can we truly trust an LLM to grade another LLM, or are we simply building a hall of mirrors?
In this session, panelists weigh all sides of model validation, testing and putting trusted agents into production.
They’ll get into the necessity of LLM-as-a-judge for real-time, automated observability at scale vs. the real world challenges of poor agent responses and the indispensable role of golden-set validation and human-in-the-loop oversight.
Attendees will walk away with discussion points and a better understanding of the need for balancing the speed of automated grading with the accuracy of human ground truth.
Format: Debate
Speakers:
- Hui Zhang
CTO and Co-Founder, Conviva - Harrison Chase
CEO, LangChain - Emmanuel Turlay
Director of engineering, CoreWeave - Sam Witteveen
Senior Technology Contributor, VentureBeat