More Assurance for Less Effort: How AI Transforms Control Testing
AI control testing automates the management testing your teams already do. It cuts the hours spent chasing evidence and raises assurance by testing every item, not just a sample.
Why it matters: less effort, more assurance
Every year, companies spend a large share of their compliance budget proving that their controls work. Most of that effort goes into chasing, collecting and reading evidence rather than judging it. AI control testing changes the equation: it automates the management testing that compliance teams and external testers already perform, so the same team can test more, finish sooner, and stand behind every conclusion with stronger evidence. It pays off in two ways.
It makes testing far more efficient. Most testing hours go to requesting, uploading and reading evidence, not to judgement. AI takes over those steps, so a test cycle that used to drag on for weeks of reminders can finish in days, and testers spend their time on the exceptions that actually need a human.
It gives you more assurance, not less. Manual testing checks a sample and hopes it is representative. AI can test every item in the population, apply exactly the same attributes every time, and cite the evidence behind each result. You find the failures a sample of 25 would have missed, and an auditor can trace every conclusion back to its source.
Adoption is still early. In a 2026 Gartner poll of 743 audit professionals, only 30% used AI for audit testing (Gartner via IT-Online). This article covers how management testing works today, how AI automates it, the tools available and the architectures that hold up in front of an auditor.
How management testing works today
Management testing is the periodic check, by a compliance team or an external tester, that key controls actually operated. Most programmes follow the same five steps:
- Scope is defined and announced. The testing team selects controls, period and samples, and notifies control owners.
- Evidence is requested. Control owners receive a request for each sampled instance: tickets, approvals, access lists, screenshots.
- Evidence is uploaded. Owners upload it to a repository, ideally the organisation's GRC system, where it is linked to the control and the test.
- Evidence is reviewed. Testers check each item against the test attributes to decide whether the control was performed correctly.
- The control is concluded effective or ineffective. Deficiencies become issues with owners and remediation dates.
In Swedish companies, step 3 is typically manual. Control owners export reports and take screenshots and upload them by hand, usually the day after the deadline and named something like screenshot_final_FINAL2.png. Automatic extraction from source systems is the stated goal almost everywhere, but rarely the reality. Most hours therefore go to requesting, uploading and reading evidence rather than to judgement.
AI testing automates the same process
AI control testing keeps the five steps and changes who, or what, performs them. Scoping and the conclusion stay with people; the evidence-heavy middle is where automation pays.
Steps 2 to 4 are where the hours go, and they are exactly the steps AI takes over.
Automation does not have to start with extraction. The review step is usually the largest, and AI can perform it on evidence that owners have uploaded by hand. That makes manual-upload environments, common in Sweden, a realistic starting point rather than a blocker.
Each automated step must leave a trace: prompts, model version and evidence hashes, so an auditor can re-perform the test.
What AI actually tests: examples
The best candidates are controls with clear yes-or-no attributes and plenty of instances, which is exactly where human testers get bored and start to skim.
| Control | What the AI checks | Typical evidence | Fit today |
|---|---|---|---|
| Change and deployment | Approved before deployment; approver is not the developer; testing documented; emergency changes approved afterwards | ITSM tickets, CI/CD pipeline logs | High |
| User access review | Review completed on time; every account marked for removal actually removed | IAM exports, review sign-off | High |
| Leavers | Access removed within the agreed deadline after termination | HR leaver list against AD and ERP users | High |
| Privileged access | Admin accounts approved, justified and reviewed | AD groups, PAM logs | High |
| Backups | Jobs completed; restore tests performed | Backup logs, restore test records | High |
| Manual journal entries | Entries above threshold approved by someone other than the preparer | ERP ledger, e.g. FGLEDG in M3 | High |
| Vendor assurance | SOC report obtained, in period, exceptions assessed | Uploaded SOC reports | Medium |
| Management review | Review performed with sufficient precision | Meeting minutes, review packs | Low: AI checks completeness, not judgement |
Take change management as the worked example. Today a tester samples, say, 25 changes and asks owners for the tickets. With AI, every change in the period is pulled from the ITSM tool and checked against four attributes, and the tester only looks at the handful that fail. The AI does not get bored on ticket number 400, and it does not need coffee.
Each result cites the evidence it rests on, so the tester can verify the one failure in seconds instead of re-reading the whole ticket.
Available tools
Most large GRC platforms now automate part of the review step, but several headline features are announced rather than available. Status is as of October 2026; performance figures are vendor claims.
| Tool | What it automates | Status |
|---|---|---|
| LogicGate Spark AI | Reviews evidence; returns pass, fail or incomplete with a rationale | Available |
| Optro (formerly AuditBoard) Autonomous Testing | Attribute tests, access reviews and reconciliations from raw evidence | Available |
| Hyperproof AI | Evidence collection and validation | Available |
| Vanta AI Agent | Evidence evaluation for SOC 2 and ISO 27001 | Available |
| Archer Evolv | Scoped AI agents with a named human supervisor | Available |
| ServiceNow IRM with ComplianceCow | Native test workflows; partner app turns control text into executable tests | Available |
| Workiva Automated Testing | Evidence, attribute and testing agents | Announced, no GA date |
Spotlight: Optro Autonomous Testing
If one product shows where control testing is heading, it is Optro, the platform formerly known as AuditBoard, which renamed itself in March 2026 (Internal Audit Guide). In May 2026 it acquired Midship and turned it into Autonomous Testing: agents that take raw evidence and perform attribute tests, access reviews and reconciliations end to end.
Three things make it worth a closer look:
- It tests, not just assists. The agent works through each attribute and documents its result, so the tester reviews outcomes instead of reading every ticket.
- It covers the whole cycle. An earlier Audit Agent handles risk-based sampling and evidence annotation, and Document Intelligence turns walkthrough notes into control narratives.
- It is open to other agents. An MCP server, launched in April 2026, lets external AI agents read and write GRC data under the platform's permissions.
The caveats are the usual ones. Midship's claim of automating up to 87% of SOX programme management is the vendor's own figure, and the product was built for US SOX teams. Swedish buyers should ask about EU data residency and about how well the agents handle evidence in Swedish.
Building your own is also an option: an LLM, plus a GRC system that exposes an API or MCP server, can automate review for your highest-volume controls. It is cheaper per control but puts the validation burden on you.
Possible architectures
Whatever the product, a sound AI testing architecture has the same five layers, with governance running alongside every layer that touches evidence or results.
The key design choice is that the AI proposes and a person concludes. Where evidence is still uploaded by hand, the GRC repository simply acts as the evidence layer; the rest of the architecture is unchanged.
Four patterns implement these layers, and most programmes end up combining two of them.
| Pattern | How it works | Strengths | Weaknesses | Best fit |
|---|---|---|---|---|
| Embedded platform agent | The GRC vendor's own agents test inside the platform (Optro, LogicGate, Archer, Workiva) | Fastest start; audit trail built in | Limited to vendor connectors; model is a black box; lock-in | Teams standardised on one platform |
| Deterministic rules generated by AI | A model writes the test rule or code once; the rule then runs without AI (ComplianceCow, SAP Joule) | Repeatable and re-performable; cheap at scale | Structured data only; the generated rule still needs review | ITGC, configuration, ERP transactions |
| External agent via MCP or API | Your own agent reads evidence and writes results into the GRC platform through MCP servers or APIs (ServiceNow Action Fabric, Optro MCP, OpenPages MCP) | Free choice of model; covers unusual controls | You own validation, security and change control | Teams with engineering capacity |
| Multi-agent pipeline | Separate agents collect evidence, test attributes and challenge results, under an orchestrator | Mirrors segregation of duties; a checker agent catches errors | Complex, costlier, harder to explain to auditors | Large, high-volume programmes |
On ServiceNow, a practical combination is native indicators and control tests for structured checks, AI Agent Studio or a partner app for evidence reading, and AI Control Tower for agent inventory and access.
Making the results auditor-ready
No audit standard setter has written AI-specific testing rules, so existing, technology-neutral rules apply in full. The PCAOB has no generative AI standard, but its amendments on technology-assisted analysis apply to fiscal years beginning on or after 15 December 2025 (Fieldguide). The IIA Global Internal Audit Standards likewise require due care and sufficient evidence whatever the tool.
In practice, an external auditor relying on AI-tested controls will ask four things:
- Is the data complete and accurate? Show where the population came from and how you reconciled it.
- Is the test re-performable? Keep prompts, model version, rule code and evidence hashes with each run.
- Who concluded? A named human signs off. ServiceNow's AI Response Assist is a useful pattern: the AI suggests, the user applies, and the AI is never recorded as author.
- Is the tester itself controlled? A test agent is part of the control environment. It needs change management, its own test suite, and segregation between whoever builds it and whoever approves its output.
EU rules add a layer. A testing agent is an ICT service, often from a third party, so it belongs in the DORA register of information and NIS2 supply-chain assessments, with least-privilege access and every action logged.
Getting started
Start where management testing hurts most, and widen the scope once your auditor accepts the method.
- Automate review of uploaded evidence first. It needs no integrations and attacks the largest share of testing hours.
- Pick five to ten pilot controls with clear attributes, such as access reviews, terminations and change approvals.
- Run in parallel for one cycle. Compare AI and manual results and measure false passes and false fails.
- Agree the evidence pack with your external auditor before go-live.
- Add automatic extraction control by control, as source-system connectors become available.
Avoid two common traps, both more tempting than a Friday fika: treating a vendor's accuracy figure as your own, and automating judgement-heavy controls too early.
Closing thoughts
AI control testing is management testing with the evidence work automated. Testers move from requesting and reading evidence to designing attributes, reviewing exceptions and concluding. The technology is ready for structured, high-volume controls today; what most programmes still need to build is the method that lets an auditor see exactly what the AI did and why. The reward: testers get their weeks back, and control owners might finally get through July without a single evidence request.