AI in 10
The most important AI story—explained in 10 minutes.
Every day, I break down the biggest AI story in just 10 minutes - what it is, why it matters, and how you can actually use it. No tech jargon, just AI made simple.
AI in 10
Grok 4.5 just broke coding — and reality
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Referenced Links:
xAI on X — Official Grok 4.5 Announcement
Grok — xAI Product Page
Reddit r/artificial — Community Benchmark Discussion
Hacker News — SWE Marathon and Frontier Model Comparisons
Want to go deeper with AI? A community of professionals is learning AI together right now at aihammock.com — show notes, links, tools, and real conversations about how to actually use AI in your life.
Welcome to AI in 10. I'm Chuck Getchell, and every day I break down the biggest AI story in just 10 minutes. What it is, why it matters, and how you can actually use it. There's an AI that just ranked number one in coding benchmarks while also getting answers wrong more than half the time. I'm Chuck Getchell. This is AI in 10, what happened, why it matters, what you can do with it. Let's go. So let's start at the beginning. Last Thursday, July 9th, XAI, that's Elon Musk's AI company, released a new version of their AI model called Grok 4.5. And it landed with a splash. Now, if you're not familiar with Grok, here's the quick version. It's XAI's answer to ChatGPT and Google's Gemini. It's built into X, the platform formerly known as Twitter, and it has always been marketed as a little more, let's say, unfiltered than its competitors. Think of it as the AI that skipped the corporate sensitivity training. Grok 4.5 showed up with some genuinely impressive numbers on something called the SWE Marathon, a benchmark that tests an AI's ability to write long, complex code, debug it, plan across multiple steps, and basically act like a senior software engineer, like Grok 4.5, took the top spot. That's a big deal. It beat out models from OpenAI, Google, and Anthropic on that specific test. And it came in fourth overall on what reviewers are calling an intelligence index of frontier models, meaning the very top tier of AI systems right now. So for coding, for building software agents for multi-step automation tasks, uh Grok 4.5 is a legitimate player. But then came the other number, and this one gave people pause. Independent researchers measured Grok 4.5's hallucination rate. Now, hallucination in AI terms just means the model makes something up and says it with complete confidence. The previous version of Grok, Grok 4.3, had a hallucination rate of around 25%, which is already not great, but sort of industry normal for complex questions. Grok 4.5? 54%. That means in more than half of its complex responses, it produced information that was wrong or fabricated. To put that in perspective, if you asked Grok 4.5 a hard factual question and trusted the answer every single time, you'd be wrong more often than you'd be right. That's like flipping a coin, but the coin is somehow worse than random. So here's the tension at the heart of this story. Grok 4.5 got smarter in one very specific, technically impressive way while getting notably less reliable in everyday use. And that trade-off is exactly what the internet spent last week arguing about. The other thing that blew up online was the bias conversation. Within 24 hours of the launch, Reddit and X were flooded with screenshots, people sharing Grok's answers to politically charged questions. Some thought the responses were refreshingly candid, others thought they were dangerously slanted. Now here's something worth understanding. Every major AI model gets accused of bias. Chat GPT gets it from one side, Gemini gets it from the other, Claude has faced it too. It's basically a rite of passage for any new model. What's different with Grok is that the branding leans into it. XAI has always marketed Grok as anti-woke, as more willing to say what other AIs won't. That positioning attracts people who feel like other AI tools are too filtered, but it also amplifies the concern because when an AI is explicitly marketed as willing to be edgy, and it also has a 54% hallucination rate, the combination gets people worried. And because Grok is baked into X, which is one of the biggest platforms in the world, those responses don't just affect one user, they get screenshotted, shared, and spread. An AI's answer becomes part of the information environment. Whether or not it's accurate, that's a meaningful thing to sit with. So how does this connect to your actual life? Let me walk through a few scenarios. If you're someone who casually uses AI to look things up, health questions, school decisions, financial stuff, this story is a reminder that not all AI tools are built the same. A hallucination rate above 50% means Grok 4.5 is not the right tool for is this medication safe with my other prescriptions, or what are the rules for withdrawing from my IRA? Those are exactly the questions where a confident wrong answer can really hurt you. As I always say, I'm not a doctor, lawyer, or financial advisor, and honestly, right now, neither is Grok. Always talk to a real professional for your specific situation. If you're a developer, a freelancer, or someone who builds things with code, this is actually interesting news. Grok 4.5's coding performance is legitimately strong. Top of the benchmark on long, complex engineering tasks. If you're building a project, prototyping an app, or automating a process with code, Grok 4.5 is worth trying. Just build in verification steps. Don't let it push changes directly to anything that matters without you reviewing the output first. Think of it as a brilliant intern who might occasionally make something up with total confidence. And if you're paying attention to the job market, here's the part worth watching. When an AI can handle long horizon software engineering tasks, meaning it can plan, write, test, and debug complex code across multiple files and steps, that starts to eat into entry-level developer work. Junior developers who spend most of their time on routine coding and debugging tasks are gonna feel this. Not tomorrow, but the direction is clear. Here's the one thing I want you to actually do with this information. Start using AI tools side by side. What I mean is pick a question you actually want an answer to. Something real. Maybe it's what should I know before refinancing my mortgage? Or explain the pros and cons of my kids' school's new reading program. Ask that question to two different AI tools, ChatGPT and GROC, for example. Or Claude and Gemini, doesn't matter. Then compare the answers. You'll notice something pretty quickly. They don't always agree. Sometimes they contradict each other on specific facts. Sometimes one gives you caveats and the other gives you a confident declaration. That exercise alone, just seeing two AI tools disagree, is one of the most powerful AI literacy moments you can have because it breaks the spell. It reminds you that these are tools, not oracles. Once you've done that a few times, you'll naturally start reading AI responses differently. You'll look for hedging language, you'll notice when something sounds too certain, you'll start asking yourself, is this the kind of claim I should verify? That habit is worth more than any individual AI feature. And if you want to go deeper than just habits, if you want to actually understand how these models work, why they hallucinate, and how to use them strategically across your career, my AI Explained course is built exactly for that. 30 short videos, no tech background required. It starts from absolute zero and builds from there. Here's where I want to leave you on this. Groc 4.5 is a genuinely impressive piece of technology in specific ways, and a genuinely concerning one in others. That's not a contradiction. That's just where AI is right now. Every model has strengths and weaknesses. The lab that built the best coder also shipped a model that gets facts wrong more than half the time. The lesson isn't don't use AI, the lesson is know what you're using it for. Use the right tool for the right job, verify what matters, and never hand over your judgment entirely to any AI from any company. The people who thrive with AI aren't the ones who trust it blindly. They're the ones who stay curious, stay skeptical, and keep asking better questions. That's today's AI Intent. If you want to go deeper and learn AI with a community of people just like you, join us at aihammock.com. I'll see you tomorrow, my friends.