If you have been building with Cursor, Claude Code, Codex, or another AI coding agent for a few months, go look at your biggest file. Not the one you think is biggest. Run the numbers.
git ls-files | grep -E '\.(ts|tsx|js|jsx)$' | xargs wc -l | grep -v ' total$' | sort -rn | head -20
I did this on a large, complex web app recently and the top file was 5,541 lines. Not the top three. The top one. There were seven files over 3,000 lines and 76 over 1,000. And this was a codebase with a mostly healthy distribution: two-thirds of its files were under 300 lines. The problem was the tail.
I don't think the people working in that codebase decided to write a 5,000-line file. I think it happened one "add this to the batch worker" at a time, and I think the agent did most of the adding.
Why AI coding agents keep making your files bigger
Here is my working theory, and I'd be interested if yours is different.
When you ask an agent to add something to a feature, it goes to the file where that feature lives and adds it there. That is the path of least resistance, and it is usually what you asked for. Adding to a file and making a new file are different jobs with different costs. Adding is one diff in one place. Making a new file means choosing a name, choosing a folder, writing an export, updating the imports, and maybe touching an index. None of that was in your request, so the agent doesn't do it.
Each individual decision is reasonable. If the file is already 2,000 lines, 200 more lines looks consistent with the existing pattern, and agents are very good at matching existing patterns. Nobody at any point says "should this file keep growing?" because that question isn't part of "add retry logic to the Salesforce sync."
The dangerous part is that this compounds. A 400-line service file becomes 800 after a sprint of agent-assisted work, and 800 becomes 1,500 after the next one. Every addition makes the next addition slightly more likely to land in the same place, because the file now covers more ground, so more things "belong" there. The agent is doing exactly what you asked, over and over, and the file is quietly turning into something no one can hold in their head.
This isn't only an AI problem. Most codebases have a file like this. AI just lets us produce the growth so quickly that we also produce the sprawl quickly, and it rarely pauses to reorganize on its own.
Do large files hurt performance? SSR vs prerendered pages
This is the first question I get asked, and the answer depends on where the code runs. Line count is not the thing that costs anything. What the code makes a machine download, parse, and run is. A big file can hurt performance in some setups and not in others, so here are the two common ones.
Server-rendered pages (SSR). Your code runs on a server for every request, or at least every cold start. A serverless function has to load and parse its whole module graph before it can answer, so a function that imports one helper from a 4,000-line module pays to load all of it. Once the function is warm, the size of the file stops mattering, because V8 compiles functions lazily and a function runs the same speed whether it sits in one file or ten. The browser then still has to download and hydrate whatever client code the page includes, which is the next case.
Prerendered pages (static or build-time). The server cost is paid once, at build, so request-time server performance doesn't depend on your file sizes at all. Build time can grow with file size, but visitors never see that. The browser is where it hurts. A prerendered page still ships JavaScript for hydration, and every byte of it has to be downloaded, parsed, compiled, and executed on the visitor's device. That work blocks the main thread, which delays interactivity and shows up in Total Blocking Time and Interaction to Next Paint. A 2,800-line client component that lands in the initial bundle can slow the page down for visitors, and on a low-end phone the difference may be large. Dead code that is still imported ships too.
Where the effect is real and where it isn't. Splitting a file does not make code faster once it is running. It helps in these cases:
- The file is on the client and some of its code is only needed after an interaction. Moving that code into its own module lets you load it with a dynamic import.
- The file is on the server, serverless, and different routes use different parts of it. Separate modules mean each function loads only what it uses.
- Bundlers can't tell which exports are unused, so the whole file ships. Tree shaking works on module boundaries and on side-effect-free code, so one file with top-level side effects can pull everything in.
If none of those apply, for example a server-only file in a long-running process or code that is already in a lazy chunk, splitting changes almost nothing about speed. Source line count also isn't the same as shipped bytes, because minification and tree shaking change the size a lot. Measure the bundle before you claim a win.
The costs listed below are real in every setup, and most of them have nothing to do with runtime speed.
What large files cost you: merge conflicts, slow reviews, and dead code
Maintenance. This is where most of the pain lives, and it's the stuff that doesn't show up on a dashboard:
- Merge conflicts. Two people (or two agent sessions) editing the same 3,000-line file may collide, even when they're working on completely unrelated parts of it.
- Code review. In a 5,000-line file, a 60-line change shows the reviewer only those 60 lines. They may not notice an existing helper the change should have used, or a rule elsewhere in the file that the change breaks.
- Testing. It can be hard to test one responsibility in isolation when it shares module scope, imports, and setup with four others. Test files for big files are often big, slow, and skipped.
- Ownership. CODEOWNERS and blame can become useless when a file belongs to everyone because it does everything.
- Dormant code waking up. Big files tend to collect functions nobody calls anymore, old feature-flag branches, and commented-out blocks. Editing nearby can bring that code back to life. An agent looking for a helper may find a stale one that looks right and call it. It may also "fix" a dead function while refactoring, which turns it back into live code. A change to a shared constant, type, or module-level variable may reach old paths nobody remembered were still reachable. In a small file, dead code is easy to spot and delete. In a 5,000-line file it can hide for years, and the edit that revives it may not look related. My dead code audit prompt is the companion for finding it first.
- Comprehension. The one-sentence test often fails. If nobody on the team can say what the file does without the word "and," it is hard to know what's safe to change.
Tooling. Editors can get laggy. Language servers may slow down on type checking and go-to-definition. Git blame and diffs get noisy. Hot reload may rebuild more than it needs to. None of these are catastrophic on their own, but they tax every edit, and you can stop noticing because it's always been that way.
Agent-side. This is the one I just learned about now (file under "found out the hard way"), and I think it's the one people underestimate.
Can an AI coding agent really read a 5,000-line file?
Not the way you'd hope.
Most agent tools cap how much of a file they read in a single pass. In Claude Code it's about 2,000 lines by default; other tools have a token budget that works out to something similar. So when the agent "reads" a 5,000-line file, it's reading the first chunk, then grepping for what it thinks it needs, then reading a few more line ranges. It is skimming, not reading.
A file that size is roughly 50,000 to 70,000 tokens. Even when a tool does pull the whole thing into context, model attention degrades on details buried in the middle of a very long input. The agent usually sees the imports at the top and the function it was asked to change. It is far less reliable with the helper 3,000 lines away that already does what it's about to write, or the module-level flag set at line 40 that its new code just invalidated.
The result is edits that look right locally and are wrong globally: a duplicated block, a missed helper, a broken invariant, a second implementation of something that already existed in the same file. Then the next session inherits that, and the file is bigger again.
There's a cost angle too. Every time the agent re-reads a big file to make a small change, you're paying for those tokens, and burning context you wanted for the actual task. Big files make every agent session more expensive and less reliable at the same time.
How many lines is too many for one file?
There's no magic number. Here's the rough guide I use:
- Under 300 to 400 lines: comfortable for humans and agents.
- 400 to 800 lines: fine if the file has one cohesive responsibility.
- Over 1,000 lines: look closely. Dig in to see if this is doing several jobs.
- Over 2,000 lines: almost certainly worth splitting.
The better test is the one-sentence test. Can you describe what the file does in one sentence without using the word "and"? If yes, leave it alone even if it's long. If no, that's your split plan, and the "and"s tell you where the seams are.
A quick way to run the test is to read the comment at the top of the file. If your agent left a long block there that walks through what the file contains ("this module handles X, and also Y, plus the helpers for Z"), that can be a sign the file is doing too much. A file with one job rarely needs a long introduction, because its name and its first few lines already say what it is. A long header usually means someone, or something, felt the need to explain how the parts fit together, and the parts are what you would split. The same goes for banner comments in the middle of a file that mark off sections. Treat this as a hint, not proof. Some files have a long header for good reasons, such as a license, a protocol description, or a tricky algorithm. But if the header is mostly a table of contents, the file probably wants to be several files.
Split by responsibility, not by line count. A file that was cut in half at line 1,500 is two files that each do everything.
Which large files to refactor first: size times churn
When I looked at those 76 files over 1,000 lines, my first instinct was to make a list and start at the top. That's the wrong instinct.
A huge file nobody touches costs almost nothing. A merely-big file that gets edited every week is where the merge conflicts, the review fatigue, and the agent mistakes are all concentrated. So the thing to rank by is size multiplied by churn, and churn is something git can tell you precisely instead of something I have to guess at.
git log --since="6 months ago" --name-only --pretty=format: -- server client | grep -v '^$' | sort | uniq -c | sort -rn
That gives you commits per file for the window. Add distinct authors per file and you have a pretty good picture of which big files are actively in the way. In the codebase I was looking at, the reasonable target wasn't "nothing over 500 lines." It was "nothing over about 1,000 without a reason, and nothing over 2,000." That's roughly 24 files, not 1,065.
How to split a large file into smaller files by responsibility
The seams depend on what kind of file it is, and they're more predictable than you'd think:
- Route files usually contain a router, request validation, handlers, and business logic that wandered in from a service. Split into a thin router per sub-resource, validation schemas in their own file, and logic pushed down into services.
- Service files split by use case or pipeline stage: validation, calculation, persistence, notifications, or per-provider logic.
- Worker files want one file per job type, plus shared worker infrastructure.
- React pages are usually several components in one file. Extract hooks for data fetching and form state, pull out sections and dialogs, and move constants and types into their own files. The page becomes a container that composes the pieces.
- Utility files (
utils.ts,helpers.ts) split by domain: dates, money, strings, permissions. A grab bag with no theme is often a good place to start, because many of its functions stand alone. Check for call chains first, though. Utility functions often call each other, soformatMoneymay callroundCurrency, which callsparseNumber. Move the whole call chain together, or put the shared function in its own module that the others import. If you split a call chain across two modules that import each other, you create a circular import. - Shared types and constants split along the same lines as the code that uses them, so each module owns its own types instead of everything importing from one giant file.
Whatever the shape, the split has to be safe:
- Tests first. If coverage is weak, write characterization tests before moving anything.
- Move code verbatim. No "improvements" along the way, so any behavior change in the diff is a mistake you can spot.
- Keep the old path working, but only for now. A small file at the original path that re-exports from the new modules keeps existing imports from breaking. Treat it as temporary. If importers keep going through it, they still load every piece, and you lose the cold start and bundle benefits from the performance section above. Move the importers to the new modules in a follow-up, then delete it. Avoid large barrel files, meaning one index file that re-exports everything from a folder, for the same reason.
- Watch module-level state and side effects. Code that runs when a file is imported, such as registering a handler, opening a connection, or filling a cache, may run in a different order after a split, or twice if two modules each create their own copy. Keep shared state in one module and have the others import it.
- Keep framework rules intact. In Next.js, a
page.tsxorroute.tscan only export specific things, includingmetadata,generateStaticParams, and route segment config options likerevalidateandmaxDuration, so those stay in the original file. If you have turned on Cache Components in Next.js 16, thedynamic,revalidate, andfetchCacheoptions are removed and you useuse cacheinstead. A"use client"directive applies per file, so a component pulled out of a client page needs its own directive unless it is only imported from files that already have one. Moving a component that uses hooks into a file that gets imported by a server component can break the build. - Check for circular imports after every move. Two new modules that import each other can fail in ways that depend on load order. A tool like
madgeor the ESLint ruleimport/no-cyclewill catch them. In a TypeScript project, madge needstypescriptinstalled and yourtsconfig.jsonpassed to it, andimport/no-cycleneeds a TypeScript parser and resolver configured. Without that setup they may skip your.tsfiles and report nothing.
Plan before you refactor: why the AI can't change anything yet
The prompt below works the same way as my dead code audit prompt: the first three phases are read-only. The agent measures, analyzes, and reports before it's allowed to move a line. There's a stop after Phase 1 so I can confirm the shortlist, and a stop after Phase 3 so I can review the plan.
What I want back from Phase 3 is something I can actually review:
| File | Lines | Commits | Responsibilities found | Proposed split | Public API impact | Test coverage | Risk | Confidence | Recommendation |
|---|---|---|---|---|---|---|---|---|---|
x.ts |
3955 | 41 | proration math, Stripe calls, email, audit log | 4 files + re-export file | None (re-exports) | Weak | Medium | H | Split now |
y.ts |
2442 | 12 | router, token exchange per provider, session | router + per-provider modules | None | Good | Low | H | Split later |
z.ts |
4243 | 3 | one big resolution pipeline | n/a | n/a | Good | n/a | M | Leave alone |
Now I can see what the agent thinks each file is doing, where it wants to cut, and how sure it is. Only after I approve that can it start Phase 4.
The large file triage super-prompt for Cursor, Claude Code, and Codex
Paste this into your coding agent from the root of the repository. Replace the bracketed placeholders first: the directories you want it looking in, anything to exclude, and your size threshold.
# Large File Triage and Split Plan
Goal: Find files in [repo/directories] that are both large and frequently changed, then plan safe splits by responsibility. No behavior changes.
Out of scope: [generated code, vendored code, migrations, lockfiles, fixtures]
Size threshold: [800] lines. Churn window: 6 months (if you change this, change the --since values below too).
## Ground rules
- Work on a separate branch and run the full test suite first for a baseline.
- Phases 1-3 are read-only, except for writing to `.refactor-notes/`.
- Never read a file over 500 lines top to bottom. Outline it first (use grep or a parser to list top-level functions, classes, components, and exports with line numbers), then read only the specific line ranges you need.
- Analyze one file at a time. Finish its notes before starting the next.
- Write findings to the notes files as you go. At the start of each phase, re-read these ground rules and the notes instead of relying on memory.
- If you're unsure about a split, mark it "Uncertain" rather than guessing.
## Phase 1: Measure (size x churn)
0. Confirm git history is complete: `git rev-parse --is-shallow-repository` must print `false`. If it prints `true`, run `git fetch --unshallow` first, or the churn numbers below will be wrong.
1. Count lines for every source file in scope and keep those at or above the threshold (adjust the extensions for your languages):
`git ls-files [repo/directories] | grep -E '\.(ts|tsx|js|jsx)$' | xargs wc -l | grep -v ' total$' | sort -rn`
2. Get commits per file over the churn window, scoped to the same directories:
`git log --since="6 months ago" --name-only --pretty=format: -- [repo/directories] | grep -v '^$' | sort | uniq -c | sort -rn`
Ignore paths that no longer exist. If a file was renamed inside the window, add the old path's count to the new one.
3. For each remaining file, also count distinct authors:
`git log --since="6 months ago" --format='%an' -- <file> | sort -u | wc -l`
4. Rank by lines x commits. Take the top 10.
Show the ranked table (file, lines, commits, authors, score) and wait for me to confirm which files to analyze.
## Phase 2: Analyze (one file at a time)
1. Build the outline yourself: every top-level unit with its line range. Save it to `.refactor-notes/<file>-outline.md`. This is the master checklist.
2. If subagents are available, split the outline into chunks of related units, about 200-500 lines each. Do not assign one function per subagent. Give each subagent its line range, the file's import block, and a list of module-level state and shared helpers. Each subagent reads only its range plus anything it must follow, and writes to its own notes file. If subagents are not available, work through the chunks yourself, one at a time.
3. For every unit in its chunk, report: purpose, what it depends on, what depends on it, and any shared state it touches.
4. Reconcile: every unit in the outline must appear in exactly one report. Re-run any that are missing before continuing.
5. You decide the responsibility groupings and proposed split, because that requires seeing the whole file. Name each responsibility in a few words. If the file has one cohesive job, mark it "Leave alone" and say why.
6. Map dependencies: who imports this file and which exports they use, and any circular-import risk. Also note call chains (unit A calls B calls C), because units in a chain should move together or share a common module, and any code that runs on import, such as registered handlers, opened connections, or module-level caches.
7. Check which tests cover it.
8. Use recent git history (`git log --oneline -n 30 -- <file>`) to see which sections change most and which change together.
9. Note framework rules that limit the split. For example, Next.js page and route files can only export certain things (`metadata`, `generateStaticParams`, route segment config), and `"use client"` or `"use server"` directives apply per file.
10. Flag units that look unused (no references found) and any long header comment that works as a table of contents. Do not delete or edit anything. Report them so I can decide whether to remove dead code before splitting.
Common seams to look for:
- Route files: router, request validation, handlers, business logic that belongs in a service.
- Service files: pipeline stages, per-provider or per-integration logic, pure calculation vs. persistence vs. side effects.
- React pages: data-fetching hooks, form state, sections, dialogs, constants and types.
- Utility files: group by domain (dates, money, strings, permissions), after checking for call chains.
- Shared types and constants: split along the code that uses them.
## Phase 3: Report
Deliver a table with these columns:
| File | Lines | Commits | Responsibilities found | Proposed split | Public API impact | Test coverage | Risk | Confidence (H/M/L) | Recommendation |
|---|---|---|---|---|---|---|---|---|---|
Recommendation is one of: Split now, Split later, Leave alone, Uncertain.
End with a suggested order of work. Stop here and wait for review.
## Phase 4: Changes (after approval only)
- One file per PR.
- If coverage is weak, add characterization tests first (tests that capture what the code currently does, so any change in behavior while moving code shows up as a failure).
- Move code verbatim. Do not rewrite, rename, or "improve" anything while moving it.
- Keep the original file path as a thin file that re-exports from the new modules, so existing imports don't break. Treat it as temporary: list the importers to update in a follow-up PR, then delete it. If importers keep using it, they still load every piece.
- Keep shared module-level state in one module that the others import. Do not let two modules each create their own copy.
- Keep the exports and directives the framework requires (such as `metadata`, `generateStaticParams`, or `"use client"`) in the files that need them.
- Extract one responsibility at a time. Run the type checker and tests after each move.
- Check for new circular imports after each move, using `madge` or ESLint `import/no-cycle`. In TypeScript, confirm the tool actually scans `.ts` files (madge needs `typescript` and the tsconfig, `import/no-cycle` needs a TypeScript parser and resolver) before trusting a clean result.
- Stop and report if results diverge from baseline.
- When done, list the new file sizes. Nothing new should exceed the threshold.
- Finally, propose (don't add) a lint `max-lines` rule that grandfathers current offenders.
The line I care about most in there is "Never read a file over 500 lines top to bottom." It sounds backwards to tell the agent not to read the file it's analyzing, but that's the whole point. If it reads 5,000 lines in one go, the middle gets lost. If it builds an outline first and then reads specific ranges, every unit gets looked at on purpose.
Why this refactoring prompt works: outline first, notes on disk, human checkpoints
No prompt eliminates attention degradation on long inputs. This one limits how much the agent is exposed to it:
- Outline first, then ranges. The agent builds a list of every top-level unit with line numbers, and that list becomes a checklist it has to reconcile against. Nothing gets skipped because the model got tired at line 3,200.
- External memory. Findings go to a
.refactor-notes/folder on disk, not into an ever-growing context window. Each phase starts by re-reading the notes and the ground rules, not by remembering them. - One file at a time, in small bounded units of work.
- Git does the prioritizing. Churn comes from real commands, not from the model's opinion of what looks important.
- Human checkpoints. You confirm the shortlist. You review the plan. The agent doesn't get to decide that a 4,000-line file is fine, or that it isn't.
A note on subagents, since a few tools support them now. They genuinely help here because each one starts with fresh context and no inherited degraded attention. But assign them chunks of 200 to 500 related lines, not one function each. Per-function loses the cross-function context (shared state, helpers), and leaves you with dozens of lossy summaries to stitch back together. The orchestrator builds the outline, hands out the chunks, and checks that every unit came back in exactly one report. Completeness comes from that reconciliation step, not from trusting the subagents. If your tool doesn't have subagents, the prompt works the same way sequentially.
What is a characterization test, and why write one before refactoring?
Phase 4 says "if coverage is weak, add characterization tests first," and that deserves a definition, because it's the thing that makes moving code safe.
A characterization test records what the code does today, not what it should do. The term comes from Michael Feathers' Working Effectively with Legacy Code, and he describes the technique in his post on characterization testing. You call the existing function with a realistic input, look at what actually comes out, and assert on exactly that.
test("prorates upgrade mid-cycle (current behavior)", () => {
const result = calculateProration({
oldPlan: "starter",
newPlan: "pro",
daysRemaining: 10,
cycleDays: 30,
});
// Value observed by running the code, not derived from a spec
expect(result).toEqual({ amountCents: 1633, creditCents: 500 });
});
A few things worth knowing:
- They lock in bugs too, on purpose. If the proration math is wrong, the test asserts the wrong number. Fix bugs in a separate PR, and comment anything suspicious with
// current behavior, may be a bug. - Cover the seams, not everything: the public functions other files call, and any path that touches money, data writes, or external APIs.
- Snapshot tests are the cheap version for functions that return large objects.
- An agent can write these for you, and it will, happily. Review them. It will characterize garbage just as confidently as it characterizes correct behavior.
How to stop large files from growing back: a max-lines rule and a module map
Splitting the top two dozen files is a project. Keeping them split is a habit, and it needs two small pieces of infrastructure, both of which the prompt will propose for you at the end of Phase 4.
A max-lines lint rule set somewhere around 800 as a warning, with the current offenders grandfathered. Not an error. The point isn't to block anyone, it's to make growth visible at review time instead of six months later.
A short module map in CLAUDE.md or AGENTS.md listing the major modules and what belongs where. This is the thing that changes the agent's default from "add it to the nearest file" to "put it where the map says it goes." A few lines is enough. If the agent knows that server/workers/ is one file per job type, it will make a new file for a new job type, because now that's the pattern it's matching.
And one line I've started adding to my own prompts when I'm about to let an agent add a feature: "If this adds more than 100 lines to a file that's already over 800, put the new code in its own file and tell me where." (This will likely need tweaking!) It's cheap, it's explicit, and it moves the "should this file keep growing?" question to the only moment anyone was ever going to ask it.

