Generative artificial intelligence (GenAI) has become a practical tool in accounting work, particularly for writing, summarization, research synthesis, and standardizing internal documentation. For small and mid-size CPA firms, however, the same tools can create heightened risks when used without guardrails, especially around client confidentiality, cybersecurity exposure, and professional responsibility. This article offers an implementation-focused framework designed for firms with limited in-house IT/security capacity, including those that rely on a small internal team or outsourced providers. The proposed approach, the โ3-Tier GenAI Stack,โ aligns AI usage with data sensitivity and operating environment so firms can capture efficiency gains while preserving client trust. The article also provides a pragmatic adoption pathway that emphasizes acceptable-use rules, prompt discipline, human review controls, and vendor due diligence.
Small CPA firms are already seeing the appeal of GenAI in everyday work: drafting client emails, structuring internal memos, summarizing public guidance, and accelerating first drafts of workpapers and deliverables. Harvard Business Review has emphasized that GenAI can reshape knowledge work by assisting in three core areas: reducing cognitive load by automating structured tasks, boosting cognitive capacity for unstructured tasks, and improving the learning process for your job.[1]
..the core obligation has not changed. CPAs remain accountable for confidentiality, due care, and the quality of client deliverables.
In an accounting context, structured tasks can include payroll processing, bank reconciliations, and rolling forward standardized workpapers, where GenAI may be most useful when paired with automation tools that handle repeatable, rules-based steps, while GenAI supports documentation, exception handling, and draft explanations. Unstructured tasks can include evaluating accounting treatment when guidance is ambiguous, tax planning with competing constraints, and other judgment-heavy decisions where the right inputs and outputs are not fully known at the start. In these cases, GenAI is best used as a drafting and issue-framing tool, and any conclusions, citations, or quantitative outputs should be verified by a CPA against authoritative guidance.
But the core obligation has not changed. CPAs remain accountable for confidentiality, due care, and the quality of client deliverables. The AICPA Code of Professional Conduct underscores confidentiality and due professional care as enduring responsibilities for members,[2] irrespective of technology assistance.
In practice, the most common risk is the inappropriate use of GenAI by staff: pasting sensitive information into the wrong tool, storing drafts in uncontrolled locations, or relying on AI outputs without adequate verification. For example, the wrong tool may be a personal (non-firm) chat account or an unapproved third-party AI tool that requires uploading documents, and uncontrolled locations may include personal cloud storage, personal email, or saved chat histories. Verification requires confirming key facts (e.g., effective dates, thresholds, exceptions, and required disclosures) against the authoritative source (e.g., FASB ASC, Internal Revenue Code and Treasury Regulations, or applicable professional standards) before sending.
What small firms need are practical operating models that help staff choose the right tool for the right task, with clear boundaries for sensitive information.
The 3-Tier GenAI Stack for small firms
The core idea is to align GenAI usage based on two questions: how sensitive is the information, and where will it be processed? This framework is deliberately pragmatic: it works whether the firm has no dedicated IT/security staff, a small internal team, or an outsourced managed service provider.
Tier 1: Public, low-risk use
In Tier 1, the firm uses public GenAI tools (such as ChatGPT, Gemini, and Claude) only when the content is genuinely nonconfidential and would be acceptable if it appeared in a public setting. The defining feature is not the tool; it is the information boundary. If staff use these tools through a firm-approved business/enterprise account, it is important to confirm that prompts are not used for model training under the vendorโs business terms; otherwise, reserve personal/consumer accounts for truly public content only. For example, firms can confirm whether customer prompts are used for model training by reviewing the AI vendorโs Terms for Business and their Data Processing Addendum (DPA) (and related enterprise privacy/data protection documentation). Such prompts typically include first drafts of website language (e.g., service descriptions, staff bios, and FAQ wording), generic engagement checklists, or summaries of publicly available guidance and standards. The practical rule is straightforward: no client-identifiable information, no proprietary templates that reveal firm methods, and no โjust this onceโ exceptions. Tier 1 can still deliver value when staff use AI for structure, tone, and brainstorming while keeping sensitive content completely out of the prompt.
Tier 2: Enterprise SaaS environments under firm administration
Tier 2 is where many small and midsize firms will do the majority of their day-to-day AI-assisted work because it aligns with how they already operateโinside Microsoft 365, Google Workspace, or other contracted enterprise systems. The defining feature of Tier 2 is that the firm can apply administrative controls: identity management, access permissions, device policies, retention rules, and audit logs. In other words, Tier 2 moves GenAI from โa consumer appโ to โa governed business capability.โ This tier is often appropriate for drafting client communications, summarizing meeting notes, producing first-pass internal memos, and accelerating documentationโprovided the firm has done basic due diligence on how the platform handles data. The key point is that โenterpriseโ is not automatically synonymous with โsafe.โ Small firms should confirm what data is retained, whether inputs are used to train or fine-tune models, what administrators can control, and what contractual protections exist. For example, basic due diligence can include confirming tenant-level retention settings for AI chats. In this context, โtenantโ refers to the firmโs own Microsoft 365, Google Workspace, or other enterprise environment. Due diligence can also include confirming how long AI chats are retained there, whether administrators can disable risky connectors/plugins, and whether audit logs are available to support monitoring and incident response. Tier 2 is therefore best described as โhigh-value, governable AI,โ but only when the firm treats vendor terms and configuration as part of risk management rather than a box-checking exercise.
Tier 3: Local or private model deployments (highest confidentiality control)
Tier 3 is the โprivate workspaceโ for GenAIโreserved for situations where the firm wants the strongest practical control that sensitive information stays within its own environment. This tier covers locally hosted or privately deployed models used for internal policy Q&A, drafting from firm-only templates, internal knowledge base search, and workflows involving sensitive client information that the firm does not want flowing through external services. It is also where small language models can be attractive: task-specific models can be cost-effective, fast, and easier to constrain for narrow use cases, and industry forecasts anticipate growing organizational use of smaller, specialized models for operational tasks.[3] In practice, task-specific use cases can include a model configured for internal policy Q&A, workpaper roll-forward drafting, or standardized client email templates. For example, small models such as Microsoft Phi, Google Gemma, or Meta Llama (subject to licensing/firm approval) can support these workflows using firm-maintained content without sending sensitive inputs to an external service.
The caution, however, is essential: Tier 3 reduces some categories of cloud exposure, but it introduces a different set of risks, especially misconfiguration, weak access control, and patching gaps. Security researchers have documented vulnerabilities affecting local-model tooling and have also identified instances where locally hosted LLM servers were unintentionally exposed to the public internet (for example, exposed Ollama servers), which is a reminder that โlocalโ models do not mean โsecure by default.โ[4] For a small firm, Tier 3 can be a powerful option when implemented with discipline, but it must be treated like any other production system: controlled access, network restrictions, monitoring/log review, and an owner who is accountable, whether that owner is internal staff or an outsourced IT provider.
A practical adoption path that fits small firms
Small firms succeed with GenAI when they combine policy, workflow design, and supervision. A workable approach begins with a short โacceptable useโ policy that employees can actually remember and follow. The policy should define what counts as confidential information, prohibit entry of client identifiers into Tier 1 tools, and require that client-facing outputs be reviewed by a CPA before release.
From there, firms should start with a controlled pilot that prioritizes low-risk wins. One useful pattern is to begin with drafting and summarization tasks in Tier 2, while reserving Tier 1 for truly public content. The pilot should also include an explicit quality step: the firm should test whether they can supervise AI outputs without increasing errors or rework.
Once the pilot workflows are established, the firm can formalize โprompt disciplineโ as a routine control. Prompt discipline is simply the practice of minimizing sensitive data exposure. In day-to-day terms, that means using placeholders instead of names, summarizing from sanitized notes instead of full documents, and treating AI as a drafting assistant rather than a decision-maker.
Finally, as firms consider Tier 3, vendor and deployment due diligence becomes unavoidable. Even if IT is outsourced, the firm should be able to answer basic questions such as: who can access the system, where is it reachable from, how are logs reviewed, how are patches applied, and how does the firm prevent staff from downloading untrusted models? The Cisco research on exposed servers is a reminder that misconfiguration can create real exposure.[5]
Governance and disclosure: small firm leadership sets the tone
GenAI governance in a small firm is less about committees and more about leadership signaling. If partners treat AI as โsomething staff uses,โ usage will drift into informal habits. If partners treat AI as a governed capability, adoption becomes both safer and more valuable.
One governance question that is increasingly practical is whether, when, and how to discuss GenAI use with clients. The Journal of Accountancy has noted that no single federal law or professional standard universally mandates disclosure of GenAI use, but it also highlights that transparency may strengthen trust when done thoughtfully and with appropriate legal consultation.[6] For many small firms, the best approach is to be prepared: have a simple explanation of how the firm uses AI, how the firm protects client information, and how human review is maintained.
Conclusion
Small CPA firms can capture meaningful productivity gains from GenAI without turning confidentiality into a gamble. The 3-Tier GenAI Stack gives firms a simple way to match tools and environments to the sensitivity of the work, while keeping professional responsibility and cybersecurity realities in view. Tier 1 enables safe efficiency for nonconfidential content, Tier 2 supports supervised workflows inside firm-controlled platforms, and Tier 3 offers the highest control for sensitive workโprovided it is deployed securely.
Cite as: “The Small Firm GenAI Stack: A Practical Playbook for Productivity and Confidentiality”, https://doi.org/10.67283/5E7844
[1] Maryam Alavi and George Westerman, โHow generative AI will transform knowledge work,โ Harvard Business Review (November 7, 2023), https://hbr.org/2023/11/how-generative-ai-will-transform-knowledge-work.
[2] American Institute of Certified Public Accountants, โAICPA code of professional conduct,โ effective December 15, 2014, updated periodically, retrieved February 18, 2026, https://pub.aicpa.org/codeofconduct/ethicsresources/et-cod.pdf.
[3] Gartner, โGartner predicts by 2027, organizations will use small, task-specific AI models three times more than general-purpose large language models,โ press release, April 9, 2025, https://www.gartner.com/en/newsroom/press-releases/2025-04-09-gartner-predicts-by-2027-organizations-will-use-small-task-specific-ai-models-three-times-more-than-general-purpose-large-language-models.
[4] Avi Lumelsky, โMore models, more ProbLLMs: New vulnerabilities in Ollamaโ (blog post), Oligo Security, October 30, 2024, https://www.oligo.security/blog/more-models-more-probllms; and Giannis Tziakouris and Elio Biasiotto, โDetecting Exposed LLM Servers: A Shodan Case Study on Ollama,โ Cisco Blogs (Security), September 1, 2025, https://blogs.cisco.com/security/detecting-exposed-llm-servers-shodan-case-study-on-ollama.
[5] Tziakouris and Elio Biasiotto.
[6] J. Michael Reese, โShould I disclose my use of gen AI to clients?,โ Journal of Accountancy, April 1, 2025, https://www.journalofaccountancy.com/issues/2025/apr/should-i-disclose-my-use-of-gen-ai-to-clients/.


