← All shows
How I AI cover art

ai

How I AI

How I AI, hosted by Claire Vo, is for anyone wondering how to actually use these magical new tools to improve the quality and efficiency of their work. In each episode, guests will share a specific, practical, and impactful way they’ve learned to use AI in their work or life. Expect 30-minute episodes, live screen sharing, and tips/tricks/workflows you can copy immediately. If you want to demystify AI and learn the skills you need to thrive in this new world, this podcast is for you.

Stories by episode

28 stories
Jul 13 · This solo builder runs 24/7 local AI on his own hardware | Alex Finn7 stories

Alex Finn Details His Extensive Local AI Hardware Setup

Alex Finn describes his current local AI hardware, which includes three Mac Studios (512GB), a DGX Spark, and a custom-built computer featuring an RTX 5090. He emphasizes that running AI models locally on his own hardware allows for unlimited, 24/7 intelligence processing, which would be prohibitively expensive with cloud-based models.

Finn's 'Software Factory' Automates Development with AI Loops

Alex Finn discusses his personal 'software factory,' which utilizes two loops of Claude code for building and reviewing tasks. Once a task is built and reviewed, a human's approval via a Slack emoji triggers a merge, streamlining the development process.

Finn's 'Red Pill Moment' for Local AI was Open Claude

Alex Finn attributes his deep dive into local AI to discovering 'Open Claude' in January. He describes an 'Aha!' moment using a Mac Mini with Open Claude, leading him to invest further in local model hardware, driven by a desire for personal AI control and the trend towards 'sovereign intelligence.'

+1 more →

Mac Studios Offer Unified Memory for Large Models, but Slow Speeds

Alex Finn explains that Mac Studios are suitable for large AI models due to their unified memory, which allows the entire system RAM to be used for processing. However, he notes that the memory bandwidth is low, resulting in very slow response times for complex models like GLM 5.2, which can take up to five minutes per prompt.

DGX Spark and Nvidia Chips for Mid-Range AI Performance

Finn contrasts Mac Studios with AI-focused computers like the DGX Spark, which offer unified memory with Nvidia (128GB) and good bandwidth via CUDA for mid-sized models like Quen 3.6. He also highlights traditional Nvidia chips, such as the RTX 5090, as the most powerful option, providing cloud-like speeds with lower VRAM but high bandwidth.

+1 more →

Older Hardware Can Still Be Useful for AI Tasks, Finn Asserts

Alex Finn suggests that even older computers like Mac Minis and laptops can still be utilized for AI tasks, such as managing memory for agents or parallelizing cloud work. He points to tools like Codex as being effective for managing these smaller tasks on less powerful hardware.

Open Source Tools Simplify Local Model Deployment

Alex Finn highlights that open-source tools like Open Claude and Hermes have significantly simplified the process of running AI models locally. Previously complex tasks of finding, configuring, and deploying models are now more accessible, allowing users to simply instruct the tool to find the best model for their hardware and use cases.

Jul 9 · GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark7 stories

OpenAI Releases GPT 5.6 Suite: Soul, Luna, and Tara

OpenAI has launched three new versions of its GPT 5.6 model: Soul, a frontier model; Tara, a balanced everyday model; and Luna, a cheaper, high-volume model. The speaker expresses a strong preference for GPT 5.6 Soul, describing it as their "heart's favorite."

GPT 5.6 Soul Priced Lower Than Fable 5

The new GPT 5.6 Soul model is more affordable than Fable 5, costing $5 per million input tokens and $30 per million output tokens, compared to Fable 5's $10 input and $50 output tokens. The speaker suspects OpenAI will continue to offer Soul within subscriptions, unlike Fable.

GPT 5.6 Achieves State-of-the-Art Performance on Benchmarks

GPT 5.6 has been evaluated as the brand new state-of-the-art model from OpenAI, achieving the highest performance on the "ultra mode" of the "How AI Vibe review benchmark" and "terminal bench 2.1". The speaker anticipates more evaluations in cybersecurity benchmarks as models become more advanced.

Speaker's Benchmark Favors GPT 5.6 Soul for Output Quality

The speaker's "Claire weighted index" for evaluating AI models, which balances LLM judging with personal taste (70% speaker, 30% LLM), found GPT 5.6 Soul to be the favorite. They praised Soul's output, particularly for prototyping and design, noting it "output the best work" and had the "highest taste score."

Fable 5 Praised for Output When Not Interacting Directly

While the speaker finds Fable 5's conversational style off-putting, likening it to an engineer unfamiliar with humans, they acknowledge its good output when direct interaction is not required. They also note that Sonnet 5 received a "gold star" for its human-like conversational ability, aside from occasional "M-dashes."

GPT 5.6 Soul Excels in Prototyping and Design Aesthetics

GPT 5.6 Soul is highlighted as the top performer for prototyping, producing functional designs that were considered the most interesting. The speaker also appreciated the design aesthetic of GPT models, likening them to "this big sun, this medium earth and this tiny moon."

OpenAI's GPT Models Show Promise in Crisp, Direct Writing

The speaker finds AI-generated writing often too recognizable and prefers a more direct style. GPT 5.6 is noted for being good at this crisp, frank writing, which the speaker appreciates. However, they also mentioned that GPT models sometimes use "M-dashes" and "slop talk."

Jun 30 · Sonnet 5 review: I ran 64 generations to find out if it's worth it11 stories

Anthropic's Claude Sonnet 5 Aims for Opus-Level Tasks at Sonnet Prices

Anthropic has released Claude Sonnet 5, with claims that it offers Opus-level task performance at Sonnet model prices. The new model is touted as being more agentic than previous Sonnet versions and is positioned as a more affordable alternative for complex tasks.

New 'How I AI' Benchmark Aims for Repeatable Model Evaluation

Tired of subjective 'vibe checks,' the podcast host is introducing the 'How I AI bench,' a new set of AI and Claude-graded benchmarks designed for repeatable testing of AI models. The benchmark will assess models on tasks like writing PRDs, solving bugs, and one-shotting designs.

Sonnet 5 Performance and Pricing Details Revealed

Anthropic's Sonnet 5 is positioned as a competitor to Opus 48 in terms of performance but at a significantly lower cost. The model is particularly highlighted for its capabilities in agentic tool use and computer work, offering a nearly comparable experience to Opus at a reduced price point.

Sonnet 5 Launch Pricing and Summer Discount

The introductory pricing for Claude Sonnet 5 is set at $2 per million input tokens and $10 per million output tokens, with this rate expected to increase slightly after the summer. The host advises interested users to test the model at its current launch prices.

Host Develops 'How I AI' Benchmark Using Claude Code

The podcast host details the creation of the 'How I AI' benchmark, built using Claude code, to provide a more repeatable and less subjective evaluation of AI models. This process involved using Claude to brainstorm benchmark design principles and tasks, focusing on builder-relevant use cases.

The 'How I AI' Benchmark Includes Human Vibe Check and LLM Scoring

The 'How I AI' benchmark incorporates a multi-stage evaluation process, including a human-scored 'vibe check' via an HTML page and LLM-based scoring for objective metrics. The process involved blind testing five models, including Sonnet 5, Opus 48, and potentially GPT 5.5, Gemini 5.2, and GLM.

Agent Personality and Voice Evaluated in Benchmark

A key component of the 'How I AI' benchmark is the evaluation of an agent's voice and personality, with the host noting a preference for Sonnet 46's conversational style. This subjective scoring aims to assess the 'vibe' of the interaction, beyond mere task completion.

Gemini 3 Pro and Sonnet 5 Tie for Top Spot in Blind Test

In a surprising turn, Gemini 3 Pro and the new Claude Sonnet 5 tied for the highest score in the blind benchmark evaluation, with GPT 4.5 also performing strongly. This outcome was unexpected, as the host noted Opus 48 and Sonnet 46 scored lower than anticipated.

Disagreement Between Human and AI Judges Highlighted

The benchmark revealed discrepancies between human subjective ratings and AI-driven evaluations, particularly concerning code quality and adherence to constraints. The host suggests that AI judges may lack the 'taste' and nuanced understanding of human evaluators.

Weighted Leaderboard Recommends Models for Specific Tasks

Following a weighted analysis balancing human opinion and backend performance, a new leaderboard was generated, with Sonnet 46 and Gemini 3 Pro leading, followed by GPT 5.5. Recommendations were made per task: GPT 5.5 for PRDs, Sonnet 46 for prototyping and conversational tasks, and Opus 48/Sonnet 5 for codebases.

Sonnet 5 Underperforms Host's Expectations in Benchmark

Despite its new release, Claude Sonnet 5 ended up at the bottom of the host's personal preference list in the 'How I AI' benchmark. The host plans to refine the benchmark for future model releases and hopes it will become an industry standard.

Jun 29 · No Figma. No Jira. No docs. How Gusto built a new product line with Claude Code | Eddie Kim (CTO)3 stories

Gusto CTO Eddie Kim Discusses 'Trash Can Method' for Rapid Product Development

Gusto CTO Eddie Kim shared how the company's new 'co-founder' product line was built by a small team of five in just 10 weeks. Kim described their approach as a 'trash can method' of software engineering, where code is frequently deleted and rebuilt due to a low cost of development, enabling rapid iteration.

Gusto CTO on Agent Development: 'It's Really Not That Scary and Complicated'

Gusto CTO Eddie Kim addressed the intimidation surrounding agent development, stating that it's essentially just an agent SDK running in the cloud. He highlighted the flexibility of switching models using an AI SDK and suggested that agent development is not as daunting as it might seem.

Gusto CTO: Co-founder Product Built Without Meetings, Specs, or Jira

Eddie Kim, CTO of Gusto, revealed that their new 'co-founder' product line was developed with a lean methodology, eschewing traditional project management tools. The team operated without meetings, tech specs, or a Jira board for tracking work, emphasizing a minimalist approach to development.