Productivity

Running Retros With Your AI Agents

Running retrospectives with AI agents to improve how you work with them, and what it reveals about your habits.

•
•9 min read
AI AgentsClaude CodeOpenClawRetrospectivesProductivity

Everyone talks about making AI agents better. Better prompts. Better models. Better tools. Almost nobody stops to ask: what is the human doing wrong?

I've been building with AI agents for months now. First with Claude Code, then with OpenClaw and a full agent team. The agents write code, review it, test it, ship it. My role has shifted from developer to product owner. I manage specs, review pull requests, and make decisions.

A couple of weeks ago, I tried something new. I asked my AI agents to run a retrospective. Not on their performance. On mine.

The question was simple:

"Based on everything we've worked on together, how can I change the way I communicate so that our interactions produce better results?"

The feedback was uncomfortably honest.

AI agents running a retrospective with a human, reviewing what went well, what didn't, and how to improve

Why a retro with an AI?

In any good team, you run retrospectives. What went well, what didn't, what do we change next time. It's one of the most effective ways to improve how a team works together.

The AI agents I work with have something most retrospective participants don't: perfect memory. Every message I sent, every spec I wrote, every time I changed direction mid-sprint. It's all there. They don't forget, they don't sugarcoat, and they don't have office politics to navigate.

So I asked them to reflect on how I worked. Not how they performed. How I performed as the human in the loop.

I did this across three different tools. The feedback was uncomfortably consistent.

Claude Code told me I don't think before I build

Claude Code is where I do most of my development work. I also use gstack's retro tool, which gives me quantitative data: how many features shipped, how many files changed, how many sessions I ran. Combined with Claude Code's own analysis of our interactions, the picture was clear.

My specs were too vague from the start. The agent noticed that for several pull requests, I gave loose initial instructions and then spent multiple rounds fine-tuning during development. The work got done, but it took longer than it should have. If I had been more precise upfront, the first attempt would have been closer to what I actually wanted.

Too many micro sessions. Instead of sitting down for focused, deep work sessions, I was jumping in and out. Quick messages, small tweaks, context lost between each one. The agent flagged this as a pattern that hurt the quality of our collaboration.

Not using QA enough. I built a whole QA agent team. I wrote about it. And then I wasn't using it consistently. The retro surfaced that I was skipping the testing step more often than I realized.

40 changes to one file in seven days. The agent pointed out that a single file had been modified 40 times in a week. That's a clear signal. The file was doing too much. I should have broken it into smaller components days earlier instead of patching the same monolith over and over.

Ship then fix, ship then fix. There was a visible pattern: I'd deploy a new feature, then immediately follow up with fixes. Sometimes multiple fixes. The agent connected the dots. I wasn't testing enough before shipping. I was using production as my test environment.

Specs rewritten mid-sprint. The agent noticed that requirements kept changing after work had already started. Things that were supposedly defined were getting rewritten halfway through. I wasn't thinking things through before kicking off implementation.

The common thread across all of this wasn't technology. The agent was performing fine. The human was the one being sloppy.

Claude Desktop sees the bigger picture

I also asked Claude Desktop for a retro. This one was different. Claude Desktop knows about my broader conversations, not just code. I use it for thinking through product decisions, brainstorming, and planning. It has a wider view of how I communicate.

I ask it to be a PM or an architect, but I never explicitly say the persona. I just start talking about product strategy and expect it to understand the role I want it to play. The result is that its responses are sometimes too generic. When I do specify the role, the output is noticeably sharper. Small change, big difference.

Voice-to-text creates ambiguity. I use Wispr Flow for most of my interactions. It's incredibly productive. I talk, it transcribes, the message goes out. But Claude Desktop noticed that some of my prompts read like stream-of-consciousness rather than clear instructions. When you write, you think while you type. When you talk, you think after you talk. The transcriptions sometimes carry that roughness. Worth knowing as a tradeoff.

The takeaway: a non-coding agent surfaces communication patterns that a coding agent never would. Different tools see different blind spots.

OpenClaw called me the bottleneck

With OpenClaw, I went with the classic retro format. What worked well? What didn't? What should we change for next time? I asked it to be honest. It was.

The feedback was blunt:

  • Too many open PRs. The agent was productive. I wasn't keeping up with reviews. Pull requests sitting open, blocking progress. This was a direct echo of something I'd already identified in a previous article. Apparently, I still hadn't fully solved it.
  • Merge faster. PRs sitting for days. Simple. Direct. No elaboration needed.
  • Be specific on design preferences upfront. When I didn't specify how something should look, the agent guessed. I rejected the guess. It tried again. Wasted cycles that a two-sentence design note would have prevented.
  • Stop mixing projects. Switching between projects in the same session created confusion about context and priorities. A source of errors and rework.
  • Use PR commands directly. I was sending messages on Telegram when I could have interacted directly on the pull request. Working closer to where the code lives would have been more efficient.

Again, the same pattern. Every piece of feedback pointed to human habits, not agent limitations.

The pattern across all three

Three different tools. Three independent retros. The same message kept coming back.

PatternWhere it showed up
Be precise upfrontAll three. Vague specs, unclear personas, loose design requirements.
Stop context-switchingClaude Code and OpenClaw. Micro sessions, project mixing, too many open PRs.
Use the tools you already haveClaude Code and OpenClaw. Built a QA team but didn't use it. Had PR commands but used Telegram.

Every time I was imprecise at the start, I paid for it with extra rounds of iteration later. Every time I scattered my attention, the quality dropped.

One agent giving you feedback might be a fluke. Three agents, independently, pointing at the same problems? That's a pattern worth taking seriously.

Retros for teams, not just individuals

Right now, I work mostly alone with my AI agents. But the principle extends far beyond a solo developer and their tools.

Imagine transcribing your standups and sending them to an AI. Not to replace the standup. To analyze it afterward. What patterns emerge over weeks? Who keeps raising the same blockers? What decisions keep getting revisited without resolution? An AI that reads four weeks of standup transcripts can surface trends that nobody in the room noticed because they were too close to it.

AI-facilitated retros for real teams. A human facilitator has biases. They have relationships with the people in the room. They might avoid hard topics. An AI agent that has observed the sprint's communication, the PR history, the ticket flow, can ask questions that a human facilitator might not. Not to replace the human element, but to supplement it with patterns that are hard to see from inside the team.

Now scale this up:

  • One person + one agent = personal habit feedback (what I've been doing).
  • One team + shared agent context = centralized learnings. Maybe three developers independently got feedback about vague specs. That's not a personal habit anymore. That's a process issue.
  • Multiple teams + aggregated retros = systemic pattern detection. If five teams surface the same problem, you've found an organizational issue. Not through surveys or management observation, but through patterns in the actual work.

Individual retros improve one person. Aggregated retros improve systems.

How to start

Pick whichever AI tool you use most. At the end of a sprint or a project phase, ask three questions:

  • What went well in our collaboration?
  • What didn't go well?
  • What should I change about how I communicate with you?

Then push for depth:

  • Be specific about the timeframe. "The last two weeks" works better than "in general."
  • Ask for quantitative evidence. How many PRs had spec changes mid-implementation? How many times did I change direction? How many files got touched repeatedly?
  • Don't argue with the feedback. The agent has no ego. It's just reporting patterns it observed. Sit with it.
  • Do it regularly. One retro is interesting. Monthly retros reveal trends. That's where the real value is.

The honest conclusion

I haven't fixed all the patterns my agents flagged. I still open too many PRs sometimes. I still rewrite specs mid-sprint when I realize I didn't think something through. Old habits don't disappear because an AI pointed them out.

But awareness changes behavior, gradually. I'm more precise in my specs now. I test before shipping more consistently. I specify the persona when I want Claude to think like a PM. Small adjustments, compounding over time.

The agents remember everything. The question is whether you're willing to ask what they noticed.