Ask a partner at a fifty-person firm where the season actually goes and you will not hear "preparing returns." You will hear about the client who sent eleven of fourteen documents, the three weeks of follow-up to get the last three, the review notes that sat for four days, and the extension that got filed because of a K-1 that arrived in September. Almost none of that is technical accounting work. Almost all of it is coordination. That gap — between where firms think the constraint is and where it actually is — determines whether an AI investment produces capacity or just produces a subscription.
What does the labor picture actually say?
The macro data is worth grounding in, because it cuts against the replacement narrative. The U.S. Bureau of Labor Statistics projects employment of accountants and auditors to grow 5 percent from 2025 to 2035, faster than the average for all occupations, with roughly 115,300 openings projected each year — many from workers retiring or changing careers. BLS also notes directly that the profession increasingly involves technology like artificial intelligence and automation, and that this shifts the work toward analytical and advisory activity rather than routine tasks.
Three things follow from that for a firm owner:
- The constraint is staffing, not demand. Firms are turning away work or extending deadlines because they cannot hire, not because clients stopped calling.
- The value of a seat is rising. If a senior associate is spending a third of the season chasing documents, that is the most expensive administrative labor in the building.
- Advisory is the growth line, and advisory requires hours that only exist if something else gives them back.
What should a firm build first?
In our engagements with professional services firms, the payback ranking for a practice between ten and two hundred people looks like this. Ordered by time to return:
- Document chase and intake completion. An agent that knows what each client owes, asks for it, follows up on a cadence, checks off what arrives, and escalates the stragglers to a human. Highest return, lowest risk, and it maps directly onto the worst part of the season.
- Client intake and qualification. New-client inquiries routed, scoped, and booked without a partner reading every email. Particularly valuable in the two months before season when the inbound spike collides with peak workload.
- Missed-call rescue and inbound coverage. Firms miss a startling volume of calls in March and April. Every one is either an existing client with an anxious question or a prospect calling three firms.
- Deadline and status communication. Proactive updates on where a file stands, which suppresses the inbound "any update?" volume that fragments preparer focus.
- Advisory prep and research synthesis. Pulling together the client picture ahead of a planning conversation. Real leverage, but it belongs behind a human review gate.
- Return preparation itself. Last, deliberately, and mostly through your existing tax software vendor rather than a custom build. This is the one everyone asks about first.
The pattern is the same one we found in AI for Law Firms: the use cases sold hardest into professional services are the ones touching billable technical judgment, and those carry the highest professional-responsibility risk with the weakest measurable return. The coordination layer around the technical work is where the capacity actually is.
How does a document-chase agent actually work?
This is worth specifying, because "AI chases your documents" is vague enough to mean nothing. A working system has five parts:
- A per-client checklist derived from last year's filed return — what was received, from whom, and when. Prior-year data is what makes this tractable; a generic checklist is noise.
- A request sequence that goes out on a schedule and adapts to what has already arrived, so clients are never asked twice for the same item.
- An intake and classification step that receives the document, identifies what it is, and marks the checklist — with a confidence threshold below which a human confirms instead.
- An escalation rule that routes to a human whenever a client asks a substantive question. "Where do I find my 1098?" is in scope. "Should I take the deduction?" is not, and the system must be structurally unable to answer it.
- A Decision Log recording every classification, every request sent, and every escalation, reviewable by the engagement partner. In a regulated profession this is not optional instrumentation.
The fourth and fifth points are the same architecture we use everywhere — Truth Boundaries constraining what the system may assert, and a Decision Log making every action auditable after the fact. The full doctrine is in Decision Log and Truth Boundaries, and the specific failure it prevents is described in AI Hallucinations in B2B. A model that invents a filing deadline or a treatment for a transaction is not a customer-service incident. It is a malpractice exposure.
What governance does a firm need before launch?
Accounting firms hold concentrated, high-sensitivity financial data under professional confidentiality obligations, which raises the governance bar relative to most industries. A workable pre-launch checklist, structured against the four core functions of the NIST AI Risk Management Framework — Govern, Map, Measure, and Manage:
- Govern. Name the partner accountable for the system. Not the IT vendor, not the software. A person with authority to switch it off.
- Map. Document exactly what client data the system touches, where it is processed, where it is retained, and whether any of it leaves your jurisdiction or your control.
- Measure. Define what "working" means numerically before launch — classification accuracy, escalation rate, false-completion rate on checklists — and review those numbers weekly during season.
- Manage. Write the rollback procedure and the incident path before go-live, not after the first bad week.
Two additional items specific to this profession. First, engagement letters and client consent language frequently predate any AI usage and should be reviewed. Second, if the firm serves clients with EU operations, the EU AI Act (Regulation 2024/1689) may bear on obligations further up the client's chain even where it does not bind the firm directly. That is a conversation to have with counsel early rather than during an engagement.
How should a firm talk about its AI to clients?
This is a real risk surface and it gets almost no attention. The FTC has been unambiguous that there is no AI exception to consumer protection law, and it has brought enforcement against firms overstating what their AI does — see Operation AI Comply. Practical guidance:
- Do not market accuracy you have not measured. "AI-verified" is a claim. If you cannot produce the measurement, do not print the word.
- Disclose where automation touches client communication. Clients react far worse to discovering it than to being told.
- Never imply AI performs the professional judgment. The signature and the responsibility belong to a licensed human, and your marketing should read that way.
What does a first-season deployment look like?
Timing matters more in this vertical than in any other we work in, because the calendar is non-negotiable. Never deploy a new system into a firm in February. The sequence that works:
- Off-season, weeks 1–3: instrument. Pull last season's real numbers — median days from engagement to complete document set, number of follow-up touches per client, extension rate and its causes. Most firms have never measured these.
- Weeks 3–6: build the chase agent against prior-year checklists for a single client segment. One segment, not the whole book.
- Weeks 6–9: shadow mode. The system drafts every request and classification; a staff member releases them. This is where the checklist logic gets corrected against reality.
- Weeks 9–12: live on the pilot segment, with partner review of the Decision Log twice weekly.
- Pre-season: expand to the full book only if the measured numbers from the pilot beat the baseline. If they do not, fix it or stop — do not scale a system that did not clear its own bar.
Firms that compress this into a six-week pre-season scramble generally end up in the pattern described in Why AI Projects Fail at $5M–$50M Businesses: a technically functional system that nobody trusts during the only ten weeks of the year when it matters.
What is the honest ceiling here?
Worth being direct about what this does not do:
- It does not replace preparers. It gives the preparers you already have a materially cleaner input and fewer interruptions.
- It does not fix a pricing problem. A firm underpricing compliance work will simply do more underpriced work faster.
- It does not survive bad data. If prior-year records are inconsistent, the checklist logic has nothing to build on, and that cleanup is the actual first project.
- It does not make judgment calls, by design. The boundary between coordination and judgment is the entire safety model, and a firm that erodes it to squeeze out more automation has traded a manageable risk for an unmanageable one.
For the cost structure behind any of these systems, see How Much Does an AI Agent Cost for a $5M–$50M Business? For the governance-first framing that should sit underneath the whole program, see How to Deploy AI Responsibly and Profitably.
The HI into AI Assessment tells you which system your firm should build first, and how many hours it gives back before next season.
Take the HI into AI Assessment →