Event Highlights

LLM-as-a-judge vs. human-in-the-loop oversight

Join our CTO and Co-Founder, Hui Zhang, at VB Transform, for the great debate on model validation.

As agents move from prototype to production, the evaluation bottleneck has become one of the biggest hurdles. Can we truly trust an LLM to grade another LLM, or are we simply building a hall of mirrors?

In this session, panelists weigh all sides of model validation, testing and putting trusted agents into production.

They’ll get into the necessity of LLM-as-a-judge for real-time, automated observability at scale vs. the real world challenges of poor agent responses and the indispensable role of golden-set validation and human-in-the-loop oversight.

Attendees will walk away with discussion points and a better understanding of the need for balancing the speed of automated grading with the accuracy of human ground truth.

Format: Debate

Speakers:

  • Hui Zhang
    CTO and Co-Founder, Conviva
  • Harrison Chase
    CEO, LangChain
  • Emmanuel Turlay
    Director of engineering, CoreWeave
  • Sam Witteveen
    Senior Technology Contributor, VentureBeat