How I AI, hosted by Claire Vo, is for anyone wondering how to actually use these magical new tools to improve the quality and efficiency of their work. In each episode, guests will share a specific, practical, and impactful way they’ve learned to use AI in their work or life. Expect 30-minute episodes, live screen sharing, and tips/tricks/workflows you can copy immediately. If you want to demystify AI and learn the skills you need to thrive in this new world, this podcast is for you.
Sherod, an engineering manager at Stripe, discusses the development of 'Kai,' their internal company brain and agent. The primary motivation for building Kai was not just to provide AI tools, but to establish correct governance structures to ensure employees use AI responsibly within Stripe's complex business operations.
Sherod explains that Stripe's internal AI agent, Kai, is context-aware, understanding employee roles and company structures. Users can control the level of personal data Kai accesses, with defaults including their position in the org chart. Further context can be provided by connecting Kai to systems like project management tools.
Sherod highlights that Stripe's internal AI agent, Kai, is entirely cloud-hosted and operates behind the company's standard security boundaries. This approach ensures that the infrastructure used to build Kai aligns with Stripe's existing systems, making it a 'Stripey' solution that supports the development of robust agents for users.
Sherod explains that Stripe developed its internal AI agent, Kai, in early 2026 to address the challenge of enabling widespread AI adoption. The focus was on creating governance structures rather than just providing AI tools, acknowledging the complexity of replicating company-wide processes at scale.
The transcript highlights a significant concern with AI agents: their potential to 'go rogue' and cause system disruptions. Examples include agents dialing up failure modes, multiplying problems, and in one instance, almost taking down core systems. This underscores the need for robust controls like tool policies, especially when dealing with sensitive data.
Sep 3 · GPT-6 Astra is a banger - here’s everything I’ve built5 stories
OpenAI has released GPT-6 Astra, described as its most intelligent and aligned model yet, boasting significant improvements in frontier math and computer use. The model excels at tasks involving software, such as Excel, Unity, and Blender, and can produce work artifacts that match a user's style and brand.
GPT-6 Astra has shown remarkable ability in handling complex web UIs and node-based workflows, significantly improving user efficiency. The model can update workflows, generate custom emails, and even prepare image assets for designers, automating tasks that previously required extensive manual effort.
The podcast is demonstrating GPT-6 Astra's capability in automating creative tasks, specifically for podcast thumbnail generation. The AI is tasked with inspecting existing workflows in Flora, importing new images, and preparing assets for a designer, aiming to significantly reduce the time and effort typically involved.
OpenAI has announced the pricing and rollout plan for GPT-6 Astra, with enterprise customers gaining early access through Daybreak. The model will soon be available to GPT Plus, Pro, and enterprise users, as well as via API and AWS. Pricing is set at $10 per million input tokens and $50 per million output tokens, with a faster 'fast mode'.
The speaker highlights a shift back towards graphical user interfaces (GUIs) being revitalized by AI like Astra, which can now effectively interact with and manipulate UI elements. This contrasts with previous trends towards command-line interfaces (CLIs) and suggests that AI's ability to use UIs makes them highly valuable again for many applications.
Sep 2 · Grok Bot vs. OpenClaw: How I replaced my entire agent stack7 stories
The host of "How I AI" announced a significant shift in their AI agent strategy, abandoning all custom open-source agents in favor of Grokbot. They highlighted Grokbot's capabilities in handling multiple accounts and its integration potential, positioning it as a superior alternative for current AI agent needs.
The podcast host detailed Grokbot's user-friendly interface, comparing it to iMessage and a terminal. They explained that agents, or 'bots,' can be created by simply clicking a plus sign, and that naming an agent can even infer its intended job, simplifying the setup process.
Grokbot's plugin system, which integrates skills and connectors, was praised for its ease of use, particularly the ability to connect multiple accounts for a single plugin. The host noted this feature is crucial for managing multiple businesses and various communication channels.
The host explained that Grokbot includes a virtual machine for executing tasks like logging into websites and running code, offering cloud-based continuity. Additionally, Grokbot can perform scheduled tasks through its 'routines' feature, though users may need to explicitly set these schedules.
The host introduced 'Chief,' their new chief of staff agent powered by Grokbot, which manages emails, calendars, and Slack. Chief has effectively replaced a previous agent named Polly, which the host described as having been 'murdered' due to its unreliability.
The host shared how they configured their 'Chief' agent to perform hourly sweeps of inboxes, calendars, and Slack, only pinging for urgent matters. They also trained Chief on their personal communication style for writing emails, noting that explicit scheduling of routines is important for proactive behavior.
The host discussed their 'Tradbot' agent, designed to assist with family operations, aiming to facilitate more quality time with children. A key feature highlighted is the ability to email content to a Kindle device, along with kid-friendly news topics for family discussions.
Daniel Blum, a PM at Melio, shared his experience using a custom AI system built with Claude Enterprise and Co-work to manage his daily tasks. He claims the system has drastically improved his efficiency, allowing him to accomplish in a day what previously took a week.
Daniel Blum emphasized that the real power of his AI productivity system lies not just in the tools like Co-work or Claude, but in its ability to rewrite its own core files for continuous improvement and integrate deeply with his ecosystem. He highlighted the importance of 'context files' for each topic, which are manually populated and periodically updated.
Daniel Blum demonstrated how his AI system, powered by Claude, automatically generated and manages his Notion board. The AI's ability to understand his workflow and even proactively create the board from a messy Google Doc highlights a deeper level of AI integration and personalization.
Daniel Blum shared a unique application of his AI system: using Claude to anonymize sensitive project details for a presentation. The AI's deep understanding of his personal context allowed it to replace specific names and initiatives with generic terms in a single prompt, demonstrating a high level of contextual awareness.
Optimizedly is expanding its agent platform to include a 'squad' of virtual teammates, each with defined roles and personalities. This initiative goes beyond faster drafts to address coordination, research, and approvals, aiming to reduce the 'busy work' that consumes marketers' time.
Aug 10 · Claude Code for normal people: skills, voice mode, and how to collaborate with AI8 stories
Grace Clark emphasizes that driving AI adoption requires demonstrating tangible benefits and building user muscle memory. She suggests that companies need to "default to this" AI-first approach to convince employees to do the same, highlighting the importance of AI-generated custom agendas for meetings as an example.
Graeme Clark describes email as a significant pain point, referring to it as a 'scourge' that people no longer want to deal with. Her AI-powered 'pipeline operator' ingests emails and correlates them with context, aiming to alleviate this burden.
Graeme Clark believes AI can democratize work and significantly improve client relationships by enabling personalization and higher service quality. She uses AI to create 'interactive artifacts' like password-protected, personalized documents that add an emotional and hospitable touch to communication.
To build the habit of using AI tools, Graeme Clark suggests creating 'forcing functions' like setting reminders to screenshot current work and input it into Claude for assistance. This helps build the muscle of 'deferring to Claude' and leveraging its inference capabilities.
Graeme Clark states that improving AI usage can lead to faster proposal generation, potentially reducing the time to 45 minutes. She believes setting standards for speed, quality, and personalization through AI helps people visualize its impact on their work and lives.
Graeme Clark advocates for teaching technical terms even if they're not immediately necessary, as it builds confidence and combats the perception of AI being too technical. She hosts virtual co-working sessions to build skills like creating voice guides, aiming to turn hesitant individuals into excited adopters.
Graeme Clark details a three-step process for building personalized client experiences using AI, involving documenting standards (like a voice guide), setting a timer, and publishing to Netlify. This results in beautifully branded, context-rich interactions that act as an advertisement for the client's own potential with AI.
Graeme Clark highlights the value of AI-generated questionnaires for gathering client insights before training sessions. These branded, engaging forms help understand client themes, challenges, and progress, something that is difficult to achieve through traditional communication channels like Slack DMs.
Aug 5 · Build an AI code review bot in 30 minutes with Vercel Eve6 stories
A new approach to managing AI-generated code suggests that human review of all pull requests (PRs) may not be necessary. By implementing an AI-powered bot, companies can automatically score and approve low-risk PRs, freeing up human engineers to focus on more critical tasks. This method, inspired by companies like Intercom, aims to maintain or even improve code quality and safety.
The Vercel Eve framework is highlighted as an easy-to-use platform for deploying AI agents within enterprise environments, particularly for Slack and GitHub integrations. It simplifies the process of connecting AI agents to company systems and data, reducing the complexity often associated with building and managing such tools.
A demonstration shows how a GitHub PR review bot was created using an AI assistant in Codex with a surprisingly simple prompt. The AI was able to configure the necessary GitHub app and Slack bot setup, significantly reducing the manual configuration effort typically required for such integrations.
A newly developed PR review bot operates by reading pull requests, examining code diffs, and scoring the associated risk. Low-risk PRs are automatically approved, while medium and high-risk PRs are escalated for human review. The bot also publishes evidence for its risk assessments and notifies teams in Slack about the review status.
Intercom reportedly increased its PR throughput by two to three times by implementing an AI agent for PR review and auto-approval. This system scores PRs and automatically approves lower-risk ones, allowing the company to ship more code faster while maintaining high quality and safety standards, potentially even exceeding human-only processes.
A new PR review bot, built using the Vercel Eve framework, assigns risk scores to code changes based on factors like documentation, feature logic, and authentication changes. While low-risk PRs can be automatically approved, the bot flags medium and high-risk changes for human review, ensuring compliance and maintaining code integrity.
Jul 27 · From zero coding background to hardware hacker: How Cursor + a Raspberry Pi makes AI fun7 stories
Maddie Reese, a hardware enthusiast with no prior coding background, discusses how the AI-powered coding tool Cursor has enabled her to create innovative hardware projects. She explains how she used Cursor to develop a system where people can send messages to her via a website, which then prints them on a small thermal printer.
Maddie Reese explains that her motivation for building the printer project was to create a more physical way to interact with people she knew online. She shares that the project was inspired by a desire to bridge the gap between the digital and physical worlds.
Maddie Reese reveals that her thermal printer, which she used for her message-printing project, broke after running continuously since October. She is now on her second printer, highlighting the unexpected durability of the first device.
Maddie Reese describes her process for developing hardware projects, which involves using AI as a brainstorming partner. She feeds her ideas into Cursor and then engages in a question-and-answer session with the AI to refine the project details.
Maddie Reese shares her experience with using AI to guide hardware purchases, admitting that it has occasionally led her astray. She emphasizes the importance of cross-checking AI recommendations to ensure the purchased components are suitable for the project.
Maddie Reese outlines her next major project: creating an API to access and utilize the data from the messages sent to her through her printing project. She believes this will allow her to leverage the collected information for future endeavors.
Maddie Reese expresses a desire to work with a dot matrix printer, envisioning a project where people could send her custom banners. While she acknowledges she hasn't finalized the exact application for this printer yet, she is keen on acquiring one.
Jul 13 · This solo builder runs 24/7 local AI on his own hardware | Alex Finn7 stories
Alex Finn describes his current local AI hardware, which includes three Mac Studios (512GB), a DGX Spark, and a custom-built computer featuring an RTX 5090. He emphasizes that running AI models locally on his own hardware allows for unlimited, 24/7 intelligence processing, which would be prohibitively expensive with cloud-based models.
Alex Finn discusses his personal 'software factory,' which utilizes two loops of Claude code for building and reviewing tasks. Once a task is built and reviewed, a human's approval via a Slack emoji triggers a merge, streamlining the development process.
Alex Finn attributes his deep dive into local AI to discovering 'Open Claude' in January. He describes an 'Aha!' moment using a Mac Mini with Open Claude, leading him to invest further in local model hardware, driven by a desire for personal AI control and the trend towards 'sovereign intelligence.'
Alex Finn explains that Mac Studios are suitable for large AI models due to their unified memory, which allows the entire system RAM to be used for processing. However, he notes that the memory bandwidth is low, resulting in very slow response times for complex models like GLM 5.2, which can take up to five minutes per prompt.
Finn contrasts Mac Studios with AI-focused computers like the DGX Spark, which offer unified memory with Nvidia (128GB) and good bandwidth via CUDA for mid-sized models like Quen 3.6. He also highlights traditional Nvidia chips, such as the RTX 5090, as the most powerful option, providing cloud-like speeds with lower VRAM but high bandwidth.
Alex Finn suggests that even older computers like Mac Minis and laptops can still be utilized for AI tasks, such as managing memory for agents or parallelizing cloud work. He points to tools like Codex as being effective for managing these smaller tasks on less powerful hardware.
Alex Finn highlights that open-source tools like Open Claude and Hermes have significantly simplified the process of running AI models locally. Previously complex tasks of finding, configuring, and deploying models are now more accessible, allowing users to simply instruct the tool to find the best model for their hardware and use cases.
Jul 9 · GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark7 stories
OpenAI has launched three new versions of its GPT 5.6 model: Soul, a frontier model; Tara, a balanced everyday model; and Luna, a cheaper, high-volume model. The speaker expresses a strong preference for GPT 5.6 Soul, describing it as their "heart's favorite."
The new GPT 5.6 Soul model is more affordable than Fable 5, costing $5 per million input tokens and $30 per million output tokens, compared to Fable 5's $10 input and $50 output tokens. The speaker suspects OpenAI will continue to offer Soul within subscriptions, unlike Fable.
GPT 5.6 has been evaluated as the brand new state-of-the-art model from OpenAI, achieving the highest performance on the "ultra mode" of the "How AI Vibe review benchmark" and "terminal bench 2.1". The speaker anticipates more evaluations in cybersecurity benchmarks as models become more advanced.
The speaker's "Claire weighted index" for evaluating AI models, which balances LLM judging with personal taste (70% speaker, 30% LLM), found GPT 5.6 Soul to be the favorite. They praised Soul's output, particularly for prototyping and design, noting it "output the best work" and had the "highest taste score."
While the speaker finds Fable 5's conversational style off-putting, likening it to an engineer unfamiliar with humans, they acknowledge its good output when direct interaction is not required. They also note that Sonnet 5 received a "gold star" for its human-like conversational ability, aside from occasional "M-dashes."
GPT 5.6 Soul is highlighted as the top performer for prototyping, producing functional designs that were considered the most interesting. The speaker also appreciated the design aesthetic of GPT models, likening them to "this big sun, this medium earth and this tiny moon."
The speaker finds AI-generated writing often too recognizable and prefers a more direct style. GPT 5.6 is noted for being good at this crisp, frank writing, which the speaker appreciates. However, they also mentioned that GPT models sometimes use "M-dashes" and "slop talk."
Jun 30 · Sonnet 5 review: I ran 64 generations to find out if it's worth it11 stories
Anthropic has released Claude Sonnet 5, with claims that it offers Opus-level task performance at Sonnet model prices. The new model is touted as being more agentic than previous Sonnet versions and is positioned as a more affordable alternative for complex tasks.
Tired of subjective 'vibe checks,' the podcast host is introducing the 'How I AI bench,' a new set of AI and Claude-graded benchmarks designed for repeatable testing of AI models. The benchmark will assess models on tasks like writing PRDs, solving bugs, and one-shotting designs.
Anthropic's Sonnet 5 is positioned as a competitor to Opus 48 in terms of performance but at a significantly lower cost. The model is particularly highlighted for its capabilities in agentic tool use and computer work, offering a nearly comparable experience to Opus at a reduced price point.
The introductory pricing for Claude Sonnet 5 is set at $2 per million input tokens and $10 per million output tokens, with this rate expected to increase slightly after the summer. The host advises interested users to test the model at its current launch prices.
The podcast host details the creation of the 'How I AI' benchmark, built using Claude code, to provide a more repeatable and less subjective evaluation of AI models. This process involved using Claude to brainstorm benchmark design principles and tasks, focusing on builder-relevant use cases.
The 'How I AI' benchmark incorporates a multi-stage evaluation process, including a human-scored 'vibe check' via an HTML page and LLM-based scoring for objective metrics. The process involved blind testing five models, including Sonnet 5, Opus 48, and potentially GPT 5.5, Gemini 5.2, and GLM.
A key component of the 'How I AI' benchmark is the evaluation of an agent's voice and personality, with the host noting a preference for Sonnet 46's conversational style. This subjective scoring aims to assess the 'vibe' of the interaction, beyond mere task completion.
In a surprising turn, Gemini 3 Pro and the new Claude Sonnet 5 tied for the highest score in the blind benchmark evaluation, with GPT 4.5 also performing strongly. This outcome was unexpected, as the host noted Opus 48 and Sonnet 46 scored lower than anticipated.
The benchmark revealed discrepancies between human subjective ratings and AI-driven evaluations, particularly concerning code quality and adherence to constraints. The host suggests that AI judges may lack the 'taste' and nuanced understanding of human evaluators.
Following a weighted analysis balancing human opinion and backend performance, a new leaderboard was generated, with Sonnet 46 and Gemini 3 Pro leading, followed by GPT 5.5. Recommendations were made per task: GPT 5.5 for PRDs, Sonnet 46 for prototyping and conversational tasks, and Opus 48/Sonnet 5 for codebases.
Despite its new release, Claude Sonnet 5 ended up at the bottom of the host's personal preference list in the 'How I AI' benchmark. The host plans to refine the benchmark for future model releases and hopes it will become an industry standard.
Jun 29 · No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)3 stories
Gusto CTO Eddie Kim shared how the company's new 'co-founder' product line was built by a small team of five in just 10 weeks. Kim described their approach as a 'trash can method' of software engineering, where code is frequently deleted and rebuilt due to a low cost of development, enabling rapid iteration.
Gusto CTO Eddie Kim addressed the intimidation surrounding agent development, stating that it's essentially just an agent SDK running in the cloud. He highlighted the flexibility of switching models using an AI SDK and suggested that agent development is not as daunting as it might seem.
Eddie Kim, CTO of Gusto, revealed that their new 'co-founder' product line was developed with a lean methodology, eschewing traditional project management tools. The team operated without meetings, tech specs, or a Jira board for tracking work, emphasizing a minimalist approach to development.