<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
  <title>Patryk Reba · Field notes</title>
  <link>https://patrykreba.com/blog</link>
  <atom:link href="https://patrykreba.com/feed.xml" rel="self" type="application/rss+xml" />
  <description>Notes from building AI products in production.</description>
  <item>
    <title>120 skills later: encoding how a repo works</title>
    <link>https://patrykreba.com/blog/skills-encode-how-you-work</link>
    <guid>https://patrykreba.com/blog/skills-encode-how-you-work</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>About 120 agent skills live across my repos. Many were installed, a few dozen were written for the repo, and those are the ones that matter. What a skill is, the kinds I keep, the gotchas they carry, and how I write one now.</description>
  </item>
  <item>
    <title>19 days from first commit to the App Store: shipping Headwind</title>
    <link>https://patrykreba.com/blog/headwind-19-days</link>
    <guid>https://patrykreba.com/blog/headwind-19-days</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Headwind is an iOS app for rehearsing hard conversations at work. It went from its first commit on 28 August to the App Store on 16 September 2026, on an engine we already had.</description>
  </item>
  <item>
    <title>A WebGL sky at 60 fps on a mid-range phone: hillclimbing this site's performance</title>
    <link>https://patrykreba.com/blog/hillclimbing-site-performance</link>
    <guid>https://patrykreba.com/blog/hillclimbing-site-performance</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>This site renders a nebula shader, a starfield and animated orbits behind every page. On a mid-range phone it pinned the main thread at 100%. Here is how I measured it, which fixes stuck, and which ones I threw away.</description>
  </item>
  <item>
    <title>A year of LoveStack: from first commit to the App Store</title>
    <link>https://patrykreba.com/blog/lovestack-first-year</link>
    <guid>https://patrykreba.com/blog/lovestack-first-year</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>LoveStack went from an empty project skeleton on 24 September 2025 to an iOS app on the App Store on 16 June 2026. This is the dated build log, and the decisions that shaped it.</description>
  </item>
  <item>
    <title>Claude Sonnet 5.5 and GPT-6.1 Sol landed in the same week</title>
    <link>https://patrykreba.com/blog/sonnet-5-5-and-gpt-6-1-sol</link>
    <guid>https://patrykreba.com/blog/sonnet-5-5-and-gpt-6-1-sol</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Anthropic shipped Claude Sonnet 5.5 on 28 September and OpenAI shipped GPT-6.1 Sol on 29 September, both at $2 in and $10 out per million tokens. What each one is, why the cost per task differs by about ten times, and where I would put each one.</description>
  </item>
  <item>
    <title>Decision records are the best prompt you'll ever write</title>
    <link>https://patrykreba.com/blog/decision-records-for-agents</link>
    <guid>https://patrykreba.com/blog/decision-records-for-agents</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>A prompt lasts one session. A decision record lasts until it is superseded, and every agent and person reads it. How I write ADRs so agents stop undoing settled decisions: titles that state the rule, invariants with tests, and pointers from the instruction file.</description>
  </item>
  <item>
    <title>Designing an MCP server people actually use: lessons from 50 tools</title>
    <link>https://patrykreba.com/blog/mcp-server-design</link>
    <guid>https://patrykreba.com/blog/mcp-server-design</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>The MCP server behind BBC Marble's ChatGPT desk has 50 tools. The design lessons that made them usable: short lists per person, names that never die, merged reads, descriptions written for the model, and refusals that say what happened.</description>
  </item>
  <item>
    <title>Designing the typed yes: human approval for AI that sends things</title>
    <link>https://patrykreba.com/blog/human-in-the-loop-approvals</link>
    <guid>https://patrykreba.com/blog/human-in-the-loop-approvals</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Drafting is cheap and sending is not. The approval patterns we use when an AI writes a company's email and WhatsApp messages: one word, bound to one version, with free undo, veto windows and no yes from an empty room.</description>
  </item>
  <item>
    <title>Evals for product prompts: judges, quotes and adversarial review</title>
    <link>https://patrykreba.com/blog/evals-for-product-prompts</link>
    <guid>https://patrykreba.com/blog/evals-for-product-prompts</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Headwind's debrief judges how a person handled a hard conversation, so every number on it has to be real. How we test product prompts: count what can be counted, verify every quote, build small eval sets from failures, and let a verifier try to refute each review finding.</description>
  </item>
  <item>
    <title>Following AI since ChatGPT: the releases that changed how I build</title>
    <link>https://patrykreba.com/blog/following-ai-since-chatgpt</link>
    <guid>https://patrykreba.com/blog/following-ai-since-chatgpt</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>I've followed every major model and tool release since ChatGPT launched, mostly without writing it down. This is the look back: which moments actually changed how I build software, and what I built at each step.</description>
  </item>
  <item>
    <title>How I read AI leaderboards: Artificial Analysis, LMArena, Epoch, METR, SWE-bench, Terminal-Bench, ARC</title>
    <link>https://patrykreba.com/blog/reading-leaderboards-together</link>
    <guid>https://patrykreba.com/blog/reading-leaderboards-together</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Seven leaderboards, seven different questions. What each one measures, who is on top of each today, why they disagree, and the checklist I use to turn them into a model choice.</description>
  </item>
  <item>
    <title>How I use Claude Code every day</title>
    <link>https://patrykreba.com/blog/how-i-use-claude-code</link>
    <guid>https://patrykreba.com/blog/how-i-use-claude-code</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>The loop I run most days across three products: write the task down, let the agent build in its own worktree, check the result with commands rather than summaries, and read the diff where a mistake would be expensive.</description>
  </item>
  <item>
    <title>I gave my portfolio a voice: building a GPT-Live twin that can show you my work</title>
    <link>https://patrykreba.com/blog/building-a-voice-twin</link>
    <guid>https://patrykreba.com/blog/building-a-voice-twin</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>The Talk to me button on this site starts a voice call with an AI version of me. It runs on GPT-Live over WebRTC, and a second, smaller model listens for requests to see something and puts it on your screen.</description>
  </item>
  <item>
    <title>Open-weight models in September 2026: Qwen3.8, GLM-5.3, Kimi K3, DeepSeek V4.1</title>
    <link>https://patrykreba.com/blog/open-weights-september-2026</link>
    <guid>https://patrykreba.com/blog/open-weights-september-2026</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Where open weights stand at the end of September: sizes, licenses, scores and prices for Qwen3.8, GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash, plus the Xiaomi model that now tops them all, and why the license file matters more than the label.</description>
  </item>
  <item>
    <title>Opus 5.5 is expensive per token and cheap per task: reading the cost data</title>
    <link>https://patrykreba.com/blog/what-a-task-costs-opus-5-5</link>
    <guid>https://patrykreba.com/blog/what-a-task-costs-opus-5-5</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Claude Opus 5.5 lists at twice the per-token price of Sonnet 5.5 and GPT-6.1 Sol. Reading Anthropic's cost post next to Artificial Analysis's per-effort data shows where it is cheap per finished task, and where it is not.</description>
  </item>
  <item>
    <title>Parallel agents without chaos: subagents and git worktrees</title>
    <link>https://patrykreba.com/blog/subagents-and-worktrees</link>
    <guid>https://patrykreba.com/blog/subagents-and-worktrees</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Subagents let one session think in parallel; git worktrees let several sessions write in parallel. How I use both, how I split and merge the work, the trap that kept my instructions out of most worktrees, and when one agent is the better answer.</description>
  </item>
  <item>
    <title>Realtime voice in production: GPT-Live over WebRTC, Gemini Live as the fallback</title>
    <link>https://patrykreba.com/blog/realtime-voice-gpt-live</link>
    <guid>https://patrykreba.com/blog/realtime-voice-gpt-live</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>How LoveStack moved voice practice to OpenAI's GPT-Live: a WebRTC call owned by the app, a Convex action that only swaps SDP, a judge model instead of in-band tags, and Gemini kept for the builds already on people's phones.</description>
  </item>
  <item>
    <title>Running a business from ChatGPT: a 50-tool MCP server</title>
    <link>https://patrykreba.com/blog/business-from-chatgpt-mcp</link>
    <guid>https://patrykreba.com/blog/business-from-chatgpt-mcp</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>How the operators of a natural stone export company in Dubai run their outreach from ChatGPT, through an MCP server that holds the record, the rules and the one door a message leaves by.</description>
  </item>
  <item>
    <title>Shipping a Capacitor iOS app with coding agents</title>
    <link>https://patrykreba.com/blog/shipping-ios-with-agents</link>
    <guid>https://patrykreba.com/blog/shipping-ios-with-agents</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>LoveStack and Headwind are React apps in a native iOS shell, and most of their commits are written with coding agents. The rules, skills and guards that let agents ship TestFlight builds without stranding the builds already on people's phones.</description>
  </item>
  <item>
    <title>Speech-to-speech went mainstream: GPT-Live, Gemini Live and what changed for builders</title>
    <link>https://patrykreba.com/blog/speech-to-speech-2026</link>
    <guid>https://patrykreba.com/blog/speech-to-speech-2026</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>OpenAI made GPT-Live generally available on 10 September and Google followed with Gemini 3.8 Live on 15 September. How the two differ, what the independent numbers say, and the five things that changed for anyone building voice.</description>
  </item>
  <item>
    <title>The AI leaderboard on 1 October 2026: what the top 12 actually tells you</title>
    <link>https://patrykreba.com/blog/intelligence-index-october-2026</link>
    <guid>https://patrykreba.com/blog/intelligence-index-october-2026</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Artificial Analysis now scores 688 models on ten evaluations. Here is the top 12 as of today, what each evaluation measures, where the labs cluster, and what a single number leaves out.</description>
  </item>
  <item>
    <title>What realtime voice AI actually costs, and how to cap it</title>
    <link>https://patrykreba.com/blog/what-voice-ai-costs</link>
    <guid>https://patrykreba.com/blog/what-voice-ai-costs</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Realtime voice bills by the minute, and a call nobody ends keeps billing. The rates we pay across LoveStack, Headwind and this site, the five layers of caps between a call and an open-ended bill, and the gaps we still have.</description>
  </item>
  <item>
    <title>Where to spend effort when an agent writes the code</title>
    <link>https://patrykreba.com/blog/spending-effort-on-agent-work</link>
    <guid>https://patrykreba.com/blog/spending-effort-on-agent-work</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>When an agent writes most of the code, my effort moves to the brief, the checks and the review, and the model's effort becomes a setting I choose per role. Effort levels, model choice, reading diffs, evals and cost.</description>
  </item>
  <item>
    <title>Which coding agent should you use in 2026?</title>
    <link>https://patrykreba.com/blog/choosing-a-coding-agent</link>
    <guid>https://patrykreba.com/blog/choosing-a-coding-agent</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>I have nine coding agents installed. Two of them do nearly all the work. What each one is for in my setup, which ones I have real evidence for, and how I would choose if I were starting today.</description>
  </item>
  <item>
    <title>Writing a CLAUDE.md agents actually follow</title>
    <link>https://patrykreba.com/blog/claude-md-that-agents-follow</link>
    <guid>https://patrykreba.com/blog/claude-md-that-agents-follow</guid>
    <pubDate>Thu, 01 Oct 2026 09:00:00 GMT</pubDate>
    <description>Patterns from the instruction files in my three repos: a map before rules, traps with their reasons, hard stops, a definition of done, links instead of copies, and the step most people skip, which is tracking the file so every worktree gets it.</description>
  </item>
</channel>
</rss>
