How AI Models Are Really Judged, with Peter Gostev (Arena / LMArena)
“A model can pass every test you write and still produce something that looks completely awful.”
Peter Gostev is head of AI capabilities at Arena (LMArena), the community-based platform where millions of real people vote in blind tests to rank AI models, born out of research at UC Berkeley. Before Arena, Peter was Head of AI at Moonpig and built a large following sharing hands-on explorations of what the latest models can actually do. He joins Georgie Healy from London for a genuinely nerdy, insider look at how models are judged and where the frontier is heading.
In this episode, Peter explains the difference between static benchmarks and human judgment, and why a model can pass every test you write and still produce something that looks completely awful. He breaks down the current state of the leaderboards, why Anthropic's models are dominating and how that tracks with real world adoption, and gives a sharp comparison of the top Western models, including why Anthropic's non-reasoning models are exceptional while OpenAI's strength lies in deep reasoning. Georgie and Peter get into why people aren't using Chinese models more despite their quality, the economics behind AI pricing and how enterprise usage is priced very differently from consumer subscriptions, why release cadence matters as much as capability, and what the wave of data centre investment means for the models arriving next. Along the way there's a fond detour on Opus 3 as the model you could talk to for hours, and why better models can sometimes feel worse.
Tune in for a clear-eyed, hype-free guide to how AI models are really evaluated, straight from someone who watches the charts move in real time.
Transcript Synced · click any line to jump ▾
Peter Gostev: Whenever models come out, quite often you see these benchmarks listed within numbers next to them. Vast majority of those are kind of static benchmarks. And what I mean by that is that someone who is expert in the area design a set of questions and they're quite often verifiable questions. So you have to have whatever the field of mathematics that they're talking about, the answer has to be an integer. That's the kind of constraint, the limitation that they have to have. Otherwise it's impossible. But if it's something like, let's say you're making a landing page, how could you possibly come up with anything meaningful there that you can actually measure?
Peter Gostev: And I come against this all the time that in my personal work, quite often I would try to come up with some tests that they can run and validate and it would then pass all the tests and what comes out is complete nonsense. It's like it looks completely awful. Right? So. But as a human who can look at it like you can instantly tell, oh, there's like a lot wrong with it. I want to see is that people are actually testing these models themselves and they develop their own opinions about how these models perform. So that's the kind of thing that I want to make sure that, that. Sorry, my. So my, my agent started taking over my.
Georgie Healy: You've got too much AI behind the scenes, mate.
Peter Gostev: Yeah, yeah,
Georgie Healy: you guys. I back myself on knowing the who's who of who's in AI. The geniuses, the pathfinders, the people that are huge behind the scenes that you might have not heard of. And this is one of those perfect examples. Peter Gostaff is hiding in plain sight. He is the AI capability leader. Arena and tens of thousands of us have been really following his work in the LLM ranking space for years. He's at arena, which doesn't just measure benchmarks, they also measure human preferences and they do this at a huge scale that hasn't big influence on how AI models are actually adopted. Now in this episode we speak about the cautionary tale of deep seek.
Georgie Healy: We talk about Mythos benchmarks and the blunt take on whether there are geographies doing billion dollar training runs outside of the US and China or not. If you ever wanted to know if the hype around a certain model is grounded in fact or not, this episode is going to give you a real understanding, a strong basis in knowing what to look for going forward. Thank you, Peter, for being on this episode. I loved this show.
Speaker C: Found a scale faster on Deel. Set up payroll for any country in minutes. Hire anyone anywhere, get visas handled fast and get back to building. Visit deel.com day one that's-double.com day one.
Georgie Healy: Hi Peter. I'm so thrilled that after finally meeting you in person, I've got you on In the Blink of AI. Thank you so much for joining us all the way from London, England. Can you start us off with your AI hack of the week?
Peter Gostev: Yeah, sure. Thanks for having me. I would say my biggest hack is just always trying to turn your half thought about maybe you've got an idea about an app or something you want to research into the actual action. So what I mean by this is that what really helps is voice. So maybe it's like multiple hacks within
Georgie Healy: the hack hack inception.
Peter Gostev: Voice to. Yeah, exactly, yeah. Voice to Voice to text. Like I use it so much, I don't know if it's obnoxious, but yeah, when I'm walking my dog or something, I just talk to my phone a lot because you're out and about and you want to just like oh, you've got this idea, you want to do something and just send it off. So I really like doing that. I've got now codecs with the remote control. So as in you can use it from your phone. That really helps a lot. Or if you're just using whatever chatgpt Claude just like really turn use voice and just whenever you have like a little thought just send it off to AI and then come back deal with it later. But I think that really helps you go from just like vaguely sort of maybe thinking to something that you can actually do yourself later.
Georgie Healy: I'm not sure if this is too much information, but I have my best ideas in the shower, you know, like shower thoughts. I need like a waterproof like, like a, like a transcription to be like because that's where I come up with my best like ideas and content and things like that. And walks are good too. Like until that happens, a walk is a great idea. Love that. Now you work at arena AI, the leading community based platform for evaluating and testing artificial intelligence models. I've been following you since before you worked there. It's just an incredible space for most of us that may not really understand the technical behind the models and really understand from a benchmark perspective where they're strong, where they're weaker and things like that. What's like the mission of the business as you see it.
Peter Gostev: Yeah, for me the most important part is that we see how do real people use AI and how they value it. So that's, for me, that's the most important part. So a lot of the time you see these benchmarks which are very useful, the static benchmarks, but they're very narrow. Like gpqa, for example, was a very nice benchmark, did us a lot of good, but in reality it's only like multiple choice questions. They're very good questions, but they're still multiple choice. I can't remember how many there are like a couple of hundred of them, maybe 300. And they condense the whole of science into that. And there's definitely a value in that because we can track on consistent basis.
Peter Gostev: But what you can't do is to capture the whole of use cases, the whole of set of judgments that humans would have. And that's what we contribute. So if you're a user, you're a lawyer, or you, I don't know, just like to play with websites or something, you want to try and create landing pages, then we can capture your judgment, your feedback as well. And I think that's really important because it gives us that more rounded perspective of how AI performs. And for me personally, the reason why I joined arena as well is that I was a big user of Arena. For me, it was really useful to calibrate myself, my understanding to how models actually perform versus the numbers that I'm seeing. So for me, that's why I was a big fan.
Georgie Healy: Yeah, I could tell your passion for it. You were doing this as just someone sharing what they're learning about the models as you were going. And frankly, myself and many others were following along because we felt that passion. We were learning a lot from it. I found this really interesting that the company was originally created by researchers from UC Berkeley. And I'm curious if, you know, when the AI labs, the anthropics, the OpenAI's started to pay attention that, oh my gosh, we're getting benchmarked and whether they like when and how they started to interact with the company.
Peter Gostev: Yeah, I feel like I don't know exactly, but it feels like from the very beginning at least, the community really cared. So it was one of those places where early on it was quite a nice way to just see where the models are. And I remember I was writing about it, I had nothing to do personally with arena at the time. And I remember writing about it and I think like that went miniviral and like even Karpathi said, it's like the best place where I can learn about how the models perform. So it was like So I don't know if they were working with the labs back then, but it felt like at least the place where community really could go and see how the models are.
Peter Gostev: And because so many new models come out all the time and it's gotten. The pace is even higher than it used to be, it's really impossible to keep track of. So we do need these kind of places. The more the merrier in terms of how we can assess those models.
Georgie Healy: I love. The more is merrier and the community vibe is definitely there. The fact that Karpathy even commented on it, these heroes in AI clearly are paying attention, but the community matters because at the end of the day, these are the people using the models as well. Right. So you mentioned to me briefly the difference between static benchmarks and human judgment. Can you explain that to the audience that they understand the difference?
Peter Gostev: Yeah. So whenever models come out, quite often you see these kind of benchmarks listed within numbers next to them and so on. Vast majority of those are kind of static benchmarks. And what I mean by that is that someone who is maybe expert in the area or they just put money into it, they design a set of questions, and they're quite often verifiable questions, and they are quite often really good, very solid work. And one good example I have is Frontier Maths benchmark from Epoch AI. Very good organization. They put a lot of effort in working with mathematicians to come up with these questions, and that way they can really track the frontier of the progress.
Peter Gostev: The challenge that they had, and they all talk about it themselves, is that it's very hard to verify complex mathematics. So they had to come up with questions which have an answer which is just like one integer. So you have to have whatever the field of mathematics that they're talking about, the answer has to be an integer because they need to be able to easily validate. They can't have 100 researchers checking every single answer because otherwise it will be impossible. So that's the kind of constraint, the limitation they have to have. Otherwise it's impossible. And I think for maths, it's reasonable to have something like that. I think they have to make some compromises, but at least it gives them a good signal.
Peter Gostev: But if it's something like, let's say you're making a landing page or you're making, I don't know, some kind of data visualization, how could you possibly come up with anything meaningful there that you can actually measure? And I come against this all the time that in my personal work, quite often I would try to come up with some tests that they can run and validate and it would then pass all the tests and what comes out is complete nonsense. It's like it looks completely awful. So it's perfectly reasonable to say that, oh, I'm just bad at creating tests. I mean, maybe, but the point is that it's. It's hard. Right? So. But as a human who can look at it like, you can instantly tell, oh, there's like a lot wrong with it.
Speaker C: I'll say.
Peter Gostev: And that's what I like about this is that.
Georgie Healy: Yeah, sorry, I was just going to say I'll speak for myself that in Australia here, everyone is obsessed with Claude. Won't hear a bad word, doesn't matter which model, doesn't matter what we say. And I love Claude too. I have a subscription to Claude. But if I want to generate images, like, here is my fireplace, I want to redesign it. The only model that does a good job for me on the Western models front is chatgpt and I think that that is currently seen as a bit of a faux pas. Everyone has to say they love Claude the most. It's just better at that stuff.
Peter Gostev: For me personally, it's interesting you say that. I mean, to be fair to Claude, they actually don't have image generation model. But for example, between the Google one, Google also has an image generation model and there was definitely in this of or like people very opinionated about like, oh, I really, really like Nano banana models. So like, I think that's the. And we also had that kind of feedback a lot. It's like, oh, why? It's like, why is the ranking like this? I really love this model and it's completely fair because also like, that's what I want to see is that people are actually testing these models themselves and they develop their own opinions about how these models perform. So that's the kind of thing that I want to make sure that. Sorry, my. So my, my agent started taking over my.
Georgie Healy: You've got too much AI behind the scenes, mate. No, that is such a good point. And it reminds me of, you know, if you judge a bird how to swim or a fish how to fly, it's like obviously not going to perform well. You've got to judge it based on what it's trying to be good at. On that point, tell the listeners, Peter, what is the current state of play on the charts? Who's doing really well and for what reason? Give us the lowdown.
Peter Gostev: If we look at the arena leaderboard, I think it's A pretty fair representation. It kind of matches my vibes, more or less. I think we can debate some of the details, but it's interesting if we went back maybe, I don't know, six to eight months, maybe a bit more. I remember we had Claude models lower down the rankings and I remember the feedback we used to get. It's like, well, I love Sony 3.5, like why is this not higher? And it would kind of rank, I can't remember exactly, but something like top 10, top 12, like something like that. And it was kind of interesting. We were wondering, is there anything kind of wrong there? But it seemed to all check out.
Peter Gostev: It's just what community preferred in this kind of blind testing. And now the anthropic models are really, really dominating across multiple leaderboards. And I think that matches really well in terms of the real world kind of adoption that we've seen. So I think the time that anthropic models went up, the rankings on our leaderboards is also the time when they started capturing a lot of market share and growing the market quite a lot. Beyond that, I think we've got that kind of overall tax leaderboard. And I think sometimes people would wonder, oh, some models like really high there, what's going on? Like I know Muse Spark for example did really well and I think it is not a bad model.
Peter Gostev: But yes, some of the rankings, I think it's worth exploring what's going on there a bit more. So one good way of doing that is if we go into some of the other subcategories for those leaderboards and that helps us to understand a bit more. So we've got like expert leaderboards and where it really looks at the subset of the different models, the kind of prompts that we have. So in there you'll see that the rankings change and so on. So then another category that I'm quite interested in is the kind of web, our code arena. Code arena in our case here measures mostly the kind of the front end performance so people can generate a website launching page and so on and then see how it performs.
Peter Gostev: So it's not necessarily like measuring, I don't know if you've got a million line code base and how it navigates that. So we're not measuring that. So is probably undervaluing some of those skills, but it is valuing really well the kind of front end skills and things like that. And there again, Claude models do really well. Although it's interesting to see how the progression has been from like 4.6 to 4.7. 4.8 like it's not like obviously massive jumps on the leaderboards. So I think that is kind of curious and I think another trend that is interesting is models from Chinese labs such as the Quen 3.7 GLM 5.1. Also recently just came out, M3 from Minimax that seems to be a really good model.
Peter Gostev: I just tested it yesterday so that feels like there's a lot of competition of of to the frontier labs that especially from China. So that, that is definitely interesting to see.
Speaker C: Founderscale faster on Deal. Set up payroll for any country in minutes, hire anyone anywhere and get visas handled fast so you stay focused on scaling. Deal takes care of onboarding, hr, it EOR benefits and compliance so your team can grow without borders. It's why more than 40,000 fast growing companies trust Deel to move fast. Visit deal.comday1 that's-double E L.com Day 1.
Georgie Healy: I can't wait to ask you a bit more about the geographical race. But, but I'm dying to know how much do you think the charts influence behavior, if at all? Do you think people go to the charts to help their model choice? Like what I will pay for a subscription for and if you think that does happen, who are the people most likely to do that? Is it the engineers because they're quite quantitative driven and tell me about the people that use your charts the most and how it impacts their behavior.
Peter Gostev: I guess what is maybe slightly less helpful if you are looking at the overall category and then you are saying well that means I've got this niche use case. So I'm going to use the top ranking model for overall category. So for that we do have subcategories where we can see for example by the different occupation for example, or different kind of use case. So we've got like vision for example, or search or document and so on. So it helps you kind of narrow down. So I think that kind of hopefully influences a bit more directly in terms of what kind of use cases you care about. So if you're a lawyer you can go to the legal leaderboard and see which models perform better.
Peter Gostev: So that's one dimension, another one I think it does, I would say influence most directly. At least the way I can, I can see it is that when the new model launches there's a whole of momentum that the launch has in terms of how do people perceive it. And if the benchmark results are bad then I think people just kind of oh, I'm not even going to try it. So because there's so much going on and if it's ranked eighth, like who cares, right? I've got my already the one I've been using is higher so we can move on. But if it does rank highly, then maybe you give it a go. So I think there's this. I don't think it's a guarantee. I don't think benchmarks themselves just sell the model, make everyone use it.
Peter Gostev: But at least it gives them that kind of in to show that. Oh yeah, maybe you want to try it. I think we did see this with some interesting models like text to image models like GPT image 2 when that came out most recently, just completely destroyed the leaderboards. The preference for the images was so much higher than any other. I think it did give a lot of momentum to OpenAI to say like, oh yeah, that's like our image model is so much better. Go try it. So that's like one example but another to say there are a couple of other examples in that space. So we've got Microsoft, for example. They released their image model and it ranked pretty well. I think it was something like top three really.
Georgie Healy: That was not on my radar at all. That's a dark horse. Yeah, I think you're right. There is so much noise, it's hard to find signal. And if like not everyone can afford or has the time to test all the models themselves even if they wanted to, the charts give some kind of direction or some kind of filtration to which you can at least be like, okay, I am willing to experiment with this model because it did perform so well. I have to ask you of the three top western models, your Gemini models, your Claude and your ChatGPT models, can you tell us what makes each of them special and yeah. What you'd best use each for?
Peter Gostev: Yeah, I would say feels like there's slightly different strengths to them like you say. So I would say between for example Anthropic and OpenAI is probably the most direct comparison between two of them. Is maybe one thing to point out the similarities. I was chatting how quickly labs release models and anthropic and OpenAI have both accelerated a lot. It used to be the case that they would release models like if you go back a while, I guess it depends how you count what you consider a model release. But the kind of meaningful model releases were much further apart, maybe like six months or something. And now it's literally like a month and a half and you get a
Georgie Healy: new model and people are still like, I want more, I want more.
Peter Gostev: Yeah, Yeah, I do. So I would say that is similar between anthropic OpenAI that is, they are quite close in that and I think it gives them a lot of advantage. One kind of negative example of that, as we saw with Gypsy, they had a lot of hype momentum and quite rightly so when they released. So that was about a year and a half ago. Deep seq v3, Jeepseek R1 came out and it was amazing. People really loved it. And it was probably at the time they were pretty close to OpenAI and I think they were above Anthropic at the time on the rankings. So that was very impressive. But then what happened is that they didn't release anything meaningful for like a year.
Peter Gostev: They did have some releases, but they were kind of incremental. They were not really that much better and they didn't position it as such. And I think that kind of really slowed them down. And if they were releasing every two months, who knows, maybe they would have been way higher. I'm sure there are other consideration constraints that they had. I'm sure they weren't just like sitting around for no reason.
Speaker C: No.
Georgie Healy: It doesn't seem that it's lack of motivation that you're not releasing models, but it is a good point.
Peter Gostev: Yeah, I think the sense I get as well that Google is a little bit slower in these kind of big changes and it's a little bit harder to pass for Google because sometimes they do the kind of 3.1 preview and then they would like maybe update it a bit and then goes ga and it's like was that the canvas is released or not? I don't, but I guess maybe for them it does count. So I don't know, it's a bit hard to parse. So I would say that's maybe the advantage that I have. In terms of the keyboard capabilities, I would say the biggest thing that stands out to me is that quite often the anthropic models, the non reasoning models were insanely good. Like if you look at the benchmarks that our benchmark and my personal testing as well, the non reasoning models are just exceptional.
Peter Gostev: Like that's the best ones out there. When you switch on the reasoning, it feels to me like it's a bit better. It's not like a lot better. But for OpenAI it's kind of the opposite. The non reasoning models kind of bad. Not like you can use them day to day, like they're fine. But in terms of like I know you want to do some coding tasks or something a bit harder. Like they were genuinely much worse. The difference between the low reasoning to the highest reasoning was insane. So it's kind of the sense that I get is that if you want a more general, quick good answer, like solid answer, anthropic models feel quite deep as well. Anthropic model is really good at that.
Peter Gostev: If you want a model to go and dig through your code base and find this little bug and do some very detailed work. It feels to me like OpenAI models are better at that. Yeah and I think it's probably because like the OpenAI models seem to be like really they were the first ones in reasoning and they seem to have really pushed in that direction now. I think that's changing a bit now because we see the anthropic models that really pushing on that more reasoning side as well. You see them having this kind of marks reasoning come out for Opus 4.8. Now I haven't actually tried it myself but the benchmarks don't look that good for max version so maybe it's kind of a more.
Peter Gostev: So it's in like it's similar to X high I think but like consumes a lot more tokens so it's kind of interesting. I would probably treat it as more experimental. They're trying to push the frontier which is like awesome. They, they should totally do that. But I think that's the kind of at least I would say that was their kind of heritage is that the anthropic models had really really good base and maybe the reasoning wasn't that strong and then open air models didn't have that good base but the reasoning was exceptionally good. I think it feels like they're slightly converging a bit more where open air models seem to be getting better at the base model level or at least the non reasoning part and then the anthropic models like trying to push the reasoning further.
Peter Gostev: So I think that's the kind dynamic that you're seeing which I think is interesting. It's hard to know is this just a bit of matter of time and things will just align and that will be it. But for now that seems to be kind of the science. I would also say on the coding side anthropic models very noticeably better. At least the current versions that are out noticeably better at the kind of the UI work. So if you want to create front end of a website or something like that, they're a lot better. They're better at the details, they're better at the kind of visual hierarchy that they would Lay out and quite often open air models would be like, I don't know, this is like a little bit like GPT 5.4 especially.
Peter Gostev: Was I as 5.5? Like a little bit better, but there's definitely a lot of room to go. Yeah, yeah, I agree it's clunky, it's overly busy.
Georgie Healy: Yeah. Okay, well, this is a bit of a spicy question, but I have had a subscription to Lovable. Really love Lovable, it's amazing. But I also have a subscription to Claude that's getting so good at ui. Peter, should I be saving money?
Peter Gostev: I think the use cases may be are slightly different, I would guess. I don't know, it depends I guess on your usage of services like Lovable. I think if you're using Lovable or I mean it's not calling Lovable out, it's more generally these kind of replit. I like replit as well. They're really good. But the reason why I care about them is not so much that or their agent is better. I mean I'm sure it is better, like maybe a bit more usable in some cases, but for me it's much more about the overall infrastructure that they will have. So for example, that it will come with a database or you can do authentication and you can host it. Things like that, end to end makes a lot of difference.
Georgie Healy: Yeah, for sure.
Peter Gostev: And that matters so much because I do prototypes all the time of things and then it's kind of a pain to then see like oh, I want to host it, I want to share it with this person and then like what they do, it's not easy. And I think what would be interesting is like if for example, anthropic OpenAI, they just say, oh yes, we're going to provide hosting, we're going to provide authentication, we're going to provide integrations, I don't know with Stripe and like all of this, then I think it's much harder to defend. It will be like much more direct fight between the two. So yeah, then I guess the best one wins.
Georgie Healy: Nate, I got my entire personal website up and running within two hours and I love it and frankly like money well spent on Lovable. But yeah, I wonder if going forward that will get closer and closer in terms of competition. Mate, you touched upon it before. I can't believe between our briefing call, which was not even a week ago and now a model has dropped. Opus 4.8 has dropped since then and this is really mean, it only just dropped. But I'm wondering if you've done Any research yet, if you've looked into it, Any insights that you can share for the listeners?
Peter Gostev: Yeah, I would say for context, the big change for Anthropic models was opus 4.5. So what happened? I need to count the number of months that it released. Was it some like four, five months? I need to remind myself.
Georgie Healy: I can't believe you said Deep SEQ was a year and a half ago. It feels like five years ago in AI terms. So my math is broken. Yeah, we'll say four or five months.
Peter Gostev: Exactly. So what happened then? Was that Anthropic. So I think it was Google who released the model. If I remember correctly. I think it was like Gemini 3 maybe, which got a lot of hypertension. That was a good release at the time. And then Anthropic said, you know what, we've got something better. So they dropped 4.5, I think like three days later. This needs a lot of fact checking. So this is.
Georgie Healy: I remember it was like late last year because it affected my founders in my cohort at Google when Gemini 3 dropped. So, yeah, maybe five, six months ago that all went down.
Peter Gostev: Yeah, yeah. And then so they released 4.5 and it was genuinely excellent model. Just like big, big jump. And I'm gonna come back to capabilities. But the, the. Another big thing that they did is that they reduced the price. I think three times. The 4.5 was three times cheaper than 4.1 Opus. And this was a big deal because Opus previously was kind of unusable. Like it was kind of better, but not that much better. And so no one seriously used it. But 4.5 was great and it was cheaper. So it was still more expensive than Gemini, but it was much closer in price and better in quality. So I would say that just put them. That was when all of this cloth psychosis started and so on.
Peter Gostev: So I think they did exceptionally well done. And then since then, they've been iterating 4.6, 4.7, now 4.8. Now that's what I was talking about. The iteration speed. That is insane to have so much iteration between those two models. So when I was doing my testing with the kind of visual stuff, it doesn't capture everything, but it's nice just to see things side by side. Is that there was a meaningful jump. If you compare between 4.5 and 4.8, you can see a massive difference with this kind of more. Create me a website. Here's a 3D generation or SVG or something like that. Something we can visually see easily. You can definitely See like a very big jump.
Peter Gostev: Now what people are complaining about a little bit, there's a kind of vibes that are kind of a little bit off is that it looks like the models are getting very capable and I think they're starting to adjust their behavior a bit more. More. So maybe being a bit more argumentative or maybe like pushing back a bit too much and things like that where it's like it's hard to capture. Exactly. And we're gonna have some rankings come out soon for 4.8, so it'll be interesting to see how that goes. But I think it's like not. It's probably a combination of that capability wise. If you measure objective elements, I mean the subjective judgments, but measurable side by side things, you can look at it and say, yeah, it's a lot better.
Peter Gostev: But then when you use it and it starts behaving a bit weirdly and maybe it's like certain patterns are a little bit different, then it's much harder for the labs to tune it. Exactly right. Because some people want one thing, other people want another thing and it's start not being obviously better.
Georgie Healy: I completely agree. I remember poor ChatGPT was too sycophantic, it was too nice, it was too friendly. And then when they made it a little bit harsher, people were like, no, my friend is gone. I miss the old model. And it's like we can't win. Like what do you want?
Peter Gostev: Yeah, yeah. And I think it is interesting that I think Claude especially had always quite special personality. I remember Opus 3, which is funny by the way, if you go to Claude AI, they still have Opus 3 in the dropdown.
Georgie Healy: Do they? It's like our old friend that you met like on traveling or something.
Peter Gostev: And generally for me Opus 3 was also a little bit special. Like I remember having these kinds of more. I mean, not seriously. I would kind of do it more from kind of model exploration level. But it was the only model that I kind of had fun talking to. Like it actually, like you could actually talk to it for hours.
Georgie Healy: I completely agree. I remember several times just having the model late at night with like, just those, you know, those thoughts you have late at night where your brain actually feels clear. And it kind of put me at ease so many times about things that I just didn't have answers to. And yeah, it's hard not to feel fondly with earlier models for that reason.
Peter Gostev: Yeah. And I think maybe there is some kind of non random reason why that happens is that if the model is kind of Useless, then maybe it doesn't really matter if it's psychophonetic or like, maybe it's like, I know it can flip out or like do some. Do some stuff that is like a little bit harmful, but maybe not. It can't actually do all that much. But if it's like, I know a mythos and it becomes, I don't know, overly aggressive or something, ghost hugs something because they got angry, I don't know, they're making it up. But you can see how maybe slightly weird behaviors that maybe people appreciated before or at least the fact that the one constraint people appreciate it now they kind of have to clamp down on it.
Peter Gostev: This capability goes as their user base grows, that maybe just kind of annoys some users. So there could be some reason why as models get better, they actually kind of get worse in other areas.
Georgie Healy: It's funny you say this, and this is not the same analogy, but I have noticed the rhetoric around AI imagery, AI generated images has really changed. When they used to give us extra fingers and the images were horrible hallucinations, it was almost fun and hilarious. Look what AI built because it was, you know, it was not threatening. And now that AI imagery is so lifelike, I'm noticing people are like, I hate AI images because it's giving me trust issues. And so it's funny like even though the models have improved, people's feeling towards them can be a bit mixed. Okay, I'm getting off topic before we get to the rapid fire. I'm dying to add, ask you about something.
Georgie Healy: Frankly, I think myself and most people don't know a lot about which is the Chinese models. Why are people not using the Chinese models more? Is it a lack of trust? Is it a lack of ability? Is it a lack of awareness? What do you think it is?
Peter Gostev: Yeah, I think it's probably a combination. One is, I would say that it's hard to know for sure, but at least anecdotally I think people are using them like a decent amount. So for example, Cursor used Kimik 2.5 as their base model to train their model. Really startups? Yeah. Their composer 2.2.5 is based on Kimi model, which is open source. So they kind of did the licensing with them then with models like, well, I guess a bunch of open source models, startups seem to be adopting them as well and so on. So I wouldn't say that, oh, no one is using them. It's hard to know for sure, but at least it's kind of more than zero. But you're right that I think the adoption is much higher for anthropic OpenAI models.
Peter Gostev: And I think that there are probably a couple of reasons. One is that it's a kind of weird market, especially on the top end that people, especially in the kind of coding assistant market is that if you're not state of the art, like why should I use you? It's like there is a kind of Pareto curve idea that okay, maybe if it's cheaper and we've got this by the way on arena as well, is that if it's much cheaper but only slightly less capable, maybe that's a good model to use. And I think for many use cases, really a valid point because. Because if you're not processing a bunch of documents, then you can sacrifice a little bit of performance. But as long as this meets your threshold, doesn't really matter.
Peter Gostev: I'd rather take 90% cost cut or 99% cost cut and just do that way cheaper. And people used to do that a lot. That's why Flash models were popular from Google. But I think in the most valuable, most hyper market right now in coding agents, I think there is so much demand for capability that anything that is off the top, even if it's like a little bit worse, why would I use it? Why my time and my opportunity cost is so high. Why would I use something else?
Georgie Healy: I have that psychology, Peter. That is my psychology. I'm like, well, if the best is still $30 a month, just use the best. The cost differentiate actual isn't that much. Like I know there's pro models at a hundred dollars a month, I don't think I would benefit from those.
Peter Gostev: I think that the, the costs are kind of spiraling a little bit. So maybe that argument starts to wane a bit. But I mean we've seen some reporting of like, I don't know, Uber is blowing through all of their budget for coding assistance like the first quarter or something.
Georgie Healy: Oh my goodness.
Peter Gostev: So like I think the CEO said that.
Georgie Healy: Do you think that will continue for everyday consumers, do you think? Because I keep hearing people running out of rate limits and purchasing additional. But I'm in the startup ecosystem. Maybe that's a very fringe and maybe it's almost a showing off thing. I am running through all my rate limits because I'm so advanced. I don't know.
Peter Gostev: Yeah, yeah, no, I think there's a. That's why this is a very valuable market. Is that the. Maybe, maybe to step back a bit what ChatGPT was trying to do, what happened I was trying to do with ChatGPT is to say, okay, we've got 900 million users, a billion users, we can charge $20 a month. I don't know, 10% of them, and we're going to make lots of money and maybe like some percentage will pay $200. The problem with that is that there's a pretty hard cap to how many people would actually pay. Now. I paid the whole time I could. So just because I want the top tier, there's so many people who like, do not care. They just give me the simple, the free one.
Peter Gostev: Why should I pay? And honestly as well, I think that kind of makes sense. Like when I'm, I know, speaking to my parents or something, it's like, oh, just use AI for this. Like, I feel bad saying, don't you want to pay $20? It just feels like it's like, oh, should they pay? It's like it's much harder decision.
Georgie Healy: I agree. I have friends in my, like, it depends on the social media channel. On LinkedIn, everyone's paying for AI and it's like obvious. But on Instagram, I like, people are just using the free tier and I, I get it. Like, they can still figure out how to get a stain out of this clothing with the free model. Then great.
Peter Gostev: Like, yeah, and I think that's, that's the reality is that if you're going for general consumer market, it's very limited how much upside you can get. And, and part of it is that like, you don't really care about productivity. Right. You've got like a few things that, like I'm saying, just like an average person going about their business, right. They maybe have like, I know 10 use cases that they care about. They want to, I don't know, check the opening hours of this. Yeah. Edit the photo, like, help me plan my meals or something like that. So like that you can do for free. You don't need like something crazy. Right. So that is the kind of the limitation that I think OpenAI hit against.
Peter Gostev: And then what really took Anthropic up in terms of their massive revenue growth is the fact that they can differentiate the amount they charge for these models. And imagine if you went to a massive company and said, oh, here's like $20 a month. And that's what kind of copilot was doing as well with Microsoft. It's like, oh, everyone's just getting $20 a month seat. That's actually a good example. They made like good amount of money for that. But it's still kind of limited, right? Even there's this kind of power curve, I think, for the users. And then what anthropic side is all, we're going to have a premium model. And by the way, you're not allowed.
Peter Gostev: If you're part of a corporation, you do not have access to, like Claude max subscription. You cannot pay $200 a month. You pay API prices. And there is like some reporting that people were doing an analysis that for $200 a month, you probably get on the order of, I don't know, $4,000 of inference that, that you get. So it's massively subsidized. So what you, you and I can go and pay for $200 to Claude to Anthropic. Like, that's not what Uber is getting Uber for that amount, they're paying for $5,000. And by the way, the usage is probably a lot higher as well because they've got, like, real things to do. Like the. I, I can take it or leave it a little bit and be like, oh, okay, next three days I'm taking time off.
Peter Gostev: I'm not going to do any side projects. But they have actual jobs to do, right? So they use them. So their usage is probably way higher. They pay way more. So they probably pay like 20 times more, like in practice. So then you can start differentiating. Your revenue really, really grows massively. And I think that's why OpenAI kind of dropped everything. Said, oh, like, damn it, we kind of missed that point. Like, we should really go after that part. And I think that's why, I think when people say like, oh, enterprise adoption and so on, there was always enterprise adoption. I think OpenAI was the leader in that at least early on in terms of like, oh, let me like plug in my.
Peter Gostev: I've got this maybe chatbot or something, and let me plug this in. We always had that. There's not 100x growth in a year on that part. I think it's just kind of growing as people adopting. There's not that much money in it. But where there is a lot of money is that kind of coding assistant world. Because people just like burning insane amount of tokens there. And I remember when I used to run AI in my previous job, we would have spent on LLMs for these kinds of. Or we've embedded AI here and it's like using AI to do that kind of task. What's tiny? I can't remember what it was exactly, but something like, I know, like $1,000 a month, like, that's us doing like a bunch of use cases, doing, like, working hard to make them work and so on.
Peter Gostev: They're delivering real value. It's tiny, but like $1,000 a month, that's like two days of someone coding. Right? So it's a very different world.
Georgie Healy: Totally. And it just reminds me so much of like, you know, anyone that has a rate card and then changes the numbers depending on whether it's an enterprise customer or if it's like a small business and things like that. It's so funny that even the big labs are doing that as well. And very fascinating behind the scenes. I don't think anyone listening, myself included, knew that that was going on. So fascinating. Peter, are you ready for the last part of the podcast, which is the rapid fire questions?
Peter Gostev: Yeah, let's do it.
Georgie Healy: All right, let's smash these out. So, number one, we've talked about Chinese models, we've talked about the Western and models. Are there other countries that are developing foundational models? True or false? Fact or fiction? Let's dispel this right now.
Peter Gostev: Yeah, pretty much. No, there are like bits and pieces. There's like, France is doing some work with Mistral. There's like here, I think is in Canada and there are some labs in Middle east and there's like, I think some Korean ones as well. So they're definitely bits and pieces. And there. It's not the, the case that no one's doing anything. But in reality, big training runs, who's doing the billion dollar training run in, in Europe? No one's doing that. So I think that's a, that's a no, unfortunately.
Georgie Healy: Thank you for dispelling that one. And at arena, can I just pay you off to rig the scores? Peter, can I just, just slip you a great meal voucher at a nice restaurant? And then you'll put my model at the top of the charts? Is that how it works?
Peter Gostev: No, we can't do that. Generally the way it's set up, we don't even have mechanisms to do that because the way it works is that we have the votes in. We filter out anything weird where, for example, the identity of the model was revealed or like some spam votes or like some something weird stuff. We filter that out and we get the rating. So there's not like, oh, let's insert our opinion kind of parameter, get flagged pretty, pretty quick.
Georgie Healy: Okay, good to know.
Peter Gostev: Yeah. So we don't have any mechanism, so no.
Georgie Healy: All right. I won't buy you a dinner voucher to put Georgina LLM in The list. Do you think Mythos will be released? And if it does, when would you predict?
Peter Gostev: I think Mythos itself. No, but they did talk about releasing some Mythos level models in the coming. Is it months? I think they said so it's probably, like, soon. Whether it's meaningful, unclear, because it could be that Opus just gets better and better, and it kind of gets to Mythos level anyway. We're probably on that trajectory. So, yeah, it's like, I feel like my sense is that it'll probably be kind of a little bit more underwhelming than we imagine. Not because it's a bad model, but if you delay by now more months and then you trim the performance a bit, then you kind of get to our kind of what you expect Opus 4.1 to be anyway. So maybe it kind of will converge.
Georgie Healy: Well said. Okay. Have there been any moments in arena history where something major has happened and you're like, why are people not talking about this? We talked about the deep seat moment. Everyone's talking about it. Gemini 3. Everyone's talking about it. What was the moment where you're like, why are people not talking about this?
Peter Gostev: Yeah, it's a. Like, we. We definitely had, like, very big moments that were actually kind of triggered by arena or, like, in combination between arena and the company. So, like, Nano Banana moment was very big. So they kept Google, kept the Nano Banana name because it was on arena kind of went viral. That's so interesting. That was very big. Big. Yeah. So these were, like, very, very big models. But I would say that, like, for us, not. Not that it's like, no one talks about it, and I've got my own opinions about, like, which models should be hyped or under hyped. But I think what is nice about it is that we are kind of part of the community, part of the conversation.
Peter Gostev: So quite often, yeah, when big models come out, people want to go check what's on the Arena. We kind of of in it together. So I think it's quite rare that really important model comes out and no one cares. That probably doesn't happen. But there's definitely a few models, some image models. I'm like, oh, that's a really. It's really good at that. I rather like that part. And then it's like, it doesn't get any hype. So I think that does happen. But maybe in this kind of small
Georgie Healy: solar level, I'm kind of jealous that you're in the room when something massive is happening and the whole world is talking about it and looking at your charts, that must feel very energizing. Now, I know you're modest and I just have to ask for One of the 65,000 people that follow you on LinkedIn, myself included. We all do want to know is there anything that you hope that people get from you most when you post and any mission that you're really pleased to be part of just by having this huge community?
Peter Gostev: Yeah, I think it's just nice to share some thoughts or some ideas, some kind of explorations that I get to do because it's such a new world of AI. There's like almost anything you do will be new by definition and there's so many untapped corners. So it's nice to have a place where we can share things and engage and discuss and so on. So yeah, for me, the interesting part, I don't get directly anything out of it, but it's much more like, yeah, let's discuss and with what do people think about this idea or direction? And I think that is a really valuable part of our broader community now that everyone wants to learn. And so many people who got much more impressive backgrounds than me also kind of participating and engaging because it's impossible for no matter how smart, senior or what kind of background you have, you cannot follow everything.
Peter Gostev: So I think it's a good place where we can just go and learn from each other a bit. So, yeah, I definitely appreciate that about the kind of. Yeah, LinkedIn, Twitter, Rajit, Discord, all the
Georgie Healy: places you're so modest. But I do relate on a tiny level of having someone so impressive add you follow you. You're like, oh my God, what is happening? But it is really nice that, that people are getting value and you're very generous with the insights you share. Okay, second last question. Future prediction. Something you believe that might be coming, that might happen in the charts in a few months that people might not be expecting.
Peter Gostev: Yeah, I think it's hard to know what people are expecting, what people are not expecting. But the one thing I think it's maybe not a groundbreaking thing, but I think the fact that we had so much investment come in, maybe that's one thing that people not appreciating is that a lot of the models that we're getting now, they're based on the investment made couple of years ago. In terms of the. What I mean by the investment is more kind of the data centers that were built out, the GPUs that were bought, that's a investment budget line maybe a year or two ago. Now these, all of these investment budget lines, they went probably like at least 2x, probably like 4x, like in a, in a year.
Peter Gostev: So the models that we should be getting soon are based off of that much higher levels of investment. So you just don't assume that what you see now that's kind of all. We are kind of at the limit. No, that's already like the money is being spent. Like we just need for the motions to come through. It's going to be there. And as long as everything works as we've seen in terms of scaling and so on, it should just get way, way better. So I think the expectations that we should have for these models that come out are super high. So we had a lot of hype from Mythos, but I don't think Mythos would be that special. I think all of the models will be there.
Peter Gostev: So the open air models, anthropic models, hopefully others like Cursor X AI combo is interesting. Could they do something special there? That would be really cool to see. Chinese models as well. Yeah, I think it's like it's not going to slow down. Our expectations should be high.
Georgie Healy: What a strong note to finish on for this incredible episode. Peter, my last question for you is where should people go to look at the arena leaderboards? Where should people follow you? And yeah, shout out to how people can follow for more breaking news.
Peter Gostev: Yeah, sure, go to arena AI. That's a good place to play with the models to look at the leaderboards. As I said, I personally loved it before I joined. That's why I joined. I would also, yeah, you can follow Arena's accounts on X on LinkedIn if you use it there. My personal accounts also there as well. But I would also suggest do check out the YouTube channel. We put like more effort, much more effort in the last like several months. And what I do there, I just did a video yesterday in a few days before that as well. It's like 40 minute videos where we look at lots of different examples of what the latest models can do. And I personally really love doing them because it's my excuse to just be like, I'm just gonna spend six hours going through the models and generate stuff and so on. So it's personally fun for me and I hope it's fun for you to watch as well.
Georgie Healy: Listeners, if you listen to this episode and you're like, I'm dying to see exactly what Peter's talking about. Visually, that is 100% where I would point you to. Next, we'll have a link in the show notes. Thank you so much, Peter, for coming on In the Blink of AI.
Peter Gostev: Brilliant. Thanks a lot for having me.
Georgie Healy: Bye. Thank you so much for listening to In the Blink of AI. If you want to go deeper on anything we've spoken about today, I write a weekly substack called Attention Is All I Need. Yes, it's hilarious. It's a pun and essentially I go into AI rants, tech news, events I'm going to and more. It's bite sized and I hear it's awesome. The link is in the show notes below.
In The Blink of AI is produced with Day One — the podcast network for founders, investors and operators. Want a show like this for your company?
Other shows worth a listen
Episodes from across the network exploring similar themes — part of the same Day One conversation.
▶Building Tech Teams with James MacDonald
Why "Show Me Your Harness" Beats Any Coding Test: Adam Witanowski, AI Architect
11 August 2026
▶Pick My Brain with Alan Jones
What's Your Moat When AI Can Copy Your Product in 48 Hours? | Dilip Jacob from Pitchberry
8 July 2026
▶Perspective X with Pauline Fetaui
Taryn Williams: Building and Exiting 6 Companies — and the Cost No One Talks About
3 July 2026
Turn podcasting into pipeline
We're the team behind the Day One Network and Blackbird's Wild Hearts. We help founders, funds and operators build trust, authority and deal flow with a show tailored to their market.