Why does AI-generated code need a different review?
AI-generated code needs a stricter review because it arrives fast, looks confident and was written by something that cannot be asked what it meant tomorrow. AI coding agents edit many files in one run, sometimes touch things you did not mention, and describe their work in summaries that can be more optimistic than the diff. The rule is simple: review the diff, not the summary.
The good news is that agent work is easy to isolate. When each task runs in its own git worktree, like a Tallos workspace, the diff contains exactly what that agent did and nothing else.
What should I check when reviewing AI-generated code?
Use this checklist on every agent diff. Scope and tests come first because they catch the most expensive mistakes early.
| # | Check | Red flags |
|---|---|---|
| 1 | Scope: the diff does what you asked and nothing more | Unrelated files, whole-file reformatting, renamed things you did not mention |
| 2 | Correctness: logic and edge cases | Only the happy path; empty, null and error cases ignored |
| 3 | Tests: new behavior is tested and the suite passes when you run it | Deleted or skipped tests, weakened assertions, bulk snapshot updates |
| 4 | Dependencies: every new package is needed and real | Packages you do not recognize, unexpected version bumps, lockfile churn |
| 5 | Security: input, auth and secrets | Hardcoded keys, disabled checks, string-built SQL, eval, permissive CORS |
| 6 | Data: migrations and defaults | Destructive migrations, changed default values, missing rollback |
| 7 | Config and CI: build, lint, hooks | Removed CI steps, disabled lint rules, --no-verify |
| 8 | Contracts: public APIs and types | Changed signatures, removed fields, silent breaking changes |
| 9 | Performance: hot paths | Queries in loops, network calls per item, unbounded lists |
| 10 | Readability: code you will maintain | Duplicated helpers, dead code, comments that describe a different implementation |
How do I review and annotate an agent's diff in Tallos?
- 1
Open the workspace's changes
When the agent finishes, click Review changes, or open Source Control in the workspace's right sidebar. The diff compares the workspace against the branch it started from, and combines staged, unstaged and untracked files.
- 2
Size up the scope first
Scan the file tree beside the hunks and the lines-added/removed chip in Source Control. Hover the chip for a code breakdown into source, tests and generated files. A small task with a large diff is your first question to the agent.
- 3
Walk every change
Read each hunk with the checklist in mind.
F7andShift+F7jump to the next and previous change. Images get side-by-side, swipe and onion-skin diffs; HTML files can open in a preview beside the diff. - 4
Run tests and try the change
Open a terminal in the same workspace and run the tests. For UI work, open the page in the built-in browser next to the diff.
- 5
Leave notes on specific lines
Hover a line and click the + in the gutter, or press
c, or use Add Review Note (Cmd+Shift+Aon macOS,Ctrl+Shift+Aon Windows). Write what is wrong and what you expect, then save withCmd+Enter. - 6
Send all notes back in one batch
Click Send notes to and pick the agent that should revise the change, or start a new agent from the same menu. Tallos composes one prompt with every note anchored to its line.
- 7
Re-review, then keep or discard
Notes stay pinned after the revision, so you can check each fix and Resolve it. When the diff is right, commit (use Generate with AI for the message if you like), push and Create PR. If it is not worth saving, delete the workspace.
Review every agent diff with line notes the agent actually reads. Get Tallos.
How do I write review notes an AI agent can act on?
Write notes an agent can act on without guessing: point at the line, say what is wrong, and say what done looks like. Batch them, because one coherent revision pass works better than a stream of single comments that pull the agent back and forth.
Vague note
- “This looks wrong.”
- “Clean this up.”
- “Add tests.”
Actionable note
- “Returns 200 on an invalid token; it should return 401.”
- “Move this retry loop into
withRetryinapi.ts; do not duplicate it.” - “Add a test for an empty cart; expect
totalto be 0.”
How do I review when several agents tried the same task?
Review the smallest correct diff first, not the first one that finished. When two or three agents worked the same task in separate workspaces, compare scope and test results side by side, keep one, and delete the rest. Where the attempts disagree, you have found the hard part of the task.
Squads can do part of this for you. In the build + review squad, a reviewer agent checks the builder's work and sends changes back until it approves; in best of three, the leader compares attempts and marks one as Recommended. You still make the final call with Accept or Discard, and you still read the diff. More in the diff review feature.
Troubleshooting: what if the review view looks wrong?
- The diff looks stale. Click refresh in the diff toolbar; an outside
gitcommand may have landed between refreshes. - Too many files to read. Ask the agent to split the change, or discard and restart with a narrower task.
- A commit hook fails. Tallos shows the hook output. Fix it yourself or use Fix with AI, which hands the failure to an agent without skipping the hook.
- The agent ignored a note. Notes that are still unresolved go into the next batch when you send again. Make the note more specific.
Frequently asked questions
Should I trust an AI agent's summary of its changes?
No. Treat the summary as a claim and the diff as the evidence. Read every changed file and run the tests yourself.
What is the most common problem in AI-generated code?
Scope creep: changes to files or behavior you did not ask for. Check scope before anything else, then tests and dependencies.
How do I send review comments back to Claude Code or Codex in Tallos?
Leave notes on lines in the workspace's diff, click Send notes to and pick the agent. Tallos sends all notes as one prompt, each anchored to its line.
Should I fix the AI's mistakes myself or send them back?
Send them back when the fix needs more than a line or two; the agent keeps context and your notes become part of its instructions. Small typo-level fixes are faster by hand.
Does Tallos merge agent code automatically?
No. Work stays on the workspace's branch until you commit, push and open a pull request, or merge it yourself.
Can an agent review another agent's code?
Yes. The build + review squad pairs a builder with a reviewer agent that sends changes back until it approves, and you accept or discard the final result.