Produced by W2D1 Media. Work with us →
Peter Gostev

Peter Gostev

Head of AI Capabilities, Arena (LMArena) at Arena

Peter Gostev is Head of AI Capabilities at Arena (LMArena), the community platform where users vote in blind tests to rank AI models. He previously served as Head of AI at Moonpig, applying generative AI to customer service and computer vision to greeting-card discovery, and built a large following sharing hands-on tests of frontier models. Gostev began his career in strategy and analytics at Accenture before leading AI strategy at NatWest Group. He created BullshitBench, an open-source benchmark testing whether language models recognise nonsensical questions rather than confidently answering them. On Day One, he discussed how AI models are really judged.

In their own words

From How AI Models Are Really Judged, with Peter Gostev (Arena / LMArena) · In The Blink Of AI
A model can pass every test you write and still produce something that looks completely awful.
▶ How AI Models Are Really Judged, with Peter Gostev (Arena / LMArena)

Career

  1. Now Head of AI Capabilities, Arena (LMArena) · Arena
  2. Earlier · Scale-up Head of AI · Moonpig
  3. Earlier Led AI strategy · NatWest Group
  4. Earlier Strategy and analytics · Accenture

Sources: ai.engineer

Connections

Spot something wrong? Suggest a correction Updated 25 September 2026

Produced by W2D1 Media

Want a show of your own?

W2D1 Media — the team behind the Day One Network, which helped build Blackbird Ventures' Wild Hearts — builds and produces podcasts for funds, founders and operators. If you want a show built around your brand, let's talk.

Start a show →
—