Part 1
Why this note exists
Between 3 and 18 August 2026 four filings landed that have almost nothing in common on the surface and everything in common underneath. Texas paused roughly 250 to 300 pending data-center interconnections for a project-by-project audit against more than 474 GW of requested load VERIFIED F04. Wood Mackenzie analysis reported on 12 August put expected commitment at roughly 28 percent of the 1,066 GW requested nationally CITED F05. Pennsylvania signed Executive Order 2026-05 making binding grid, community and transparency commitments a precondition of a data-center permit VERIFIED F03. And the FDA published a discussion paper on generative AI-enabled medical devices with 26 questions weighted toward postmarket monitoring VERIFIED F02.
Read individually they are an energy story, a market-sizing story, a siting story and a health story. Read together they describe one shift, and it is the shift this note is about.
The shift in one sentence
The burden of proof has moved off the model and onto two things almost nobody instrumented: where the compute physically sits and what it costs the grid, and how the system behaves after it goes live.
Both are evidence problems. Neither is solved by choosing a better model, and neither is solved by a policy document. They are solved by artifacts: filings, logs, samples, thresholds, and a named person who owns each one.
What follows is the pack structure we use with clients in financial services, healthcare, manufacturing and energy. It assumes you already have agents in production or a compute footprint under development. It does not assume you have a governance platform, because most of the organizations that need this most do not have one yet and cannot wait for procurement.
Part 2
Evidence Pack A: Proof of Power
For any organization with a compute footprint under development, whether owned, colocated or contracted through a partner. Six artifacts, all obtainable from parties you already pay.
| Artifact | What good looks like | Why it is now load-bearing |
|---|---|---|
| A1. Queue position and study batch | Named interconnection queue, position, study batch identifier and the current scheduled study date. | ERCOT missed its 7 August Batch Zero deadline and sought a good-cause exception, so batch dates are moving VERIFIED F04. |
| A2. Firm generation evidence | Signed generation agreement or on-site build plan with capacity, commissioning date and counterparty. | Pennsylvania effectively requires developers to bring their own generation as a permit condition VERIFIED F03. |
| A3. Infrastructure cost allocation | Written allocation of who funds the network upgrades serving the new load, with amounts. | The Pennsylvania order requires developers to fund infrastructure serving their demand rather than socialize it VERIFIED F03. |
| A4. Local approval status | Municipal or county approval evidence, dated, plus the objections raised and how they were answered. | Local approval is a precondition for DEP permit evaluation, not a later formality VERIFIED F03. |
| A5. Audit exposure statement | Explicit written answer to whether the project sits inside a paused or audited cohort. | Roughly 250 to 300 ERCOT projects are inside a verification audit right now VERIFIED F04. |
| A6. Conversion realism note | Your own assessment of the probability this capacity converts, with the assumption written down. | Market analysis expects roughly 28 percent of requested national capacity to be committed CITED F05. A plan that assumes 100 percent is a plan with an unstated bet in it. |
The transparency change most people missed
Pennsylvania's order prohibits non-disclosure agreements for data-center projects VERIFIED F03. If your siting strategy relied on confidentiality to manage community reaction, that lever is gone in at least one large market and the direction of travel is visible. Plan the community engagement as a public process from the first meeting, because it will become one.
Efficiency is now a siting argument, not just a cost argument
When interconnection is the binding constraint, throughput per watt stops being a procurement footnote. Cerebras announced the CS-4 on 18 August 2026 claiming up to 10 times the throughput per watt of the CS-3, alongside claimed speed advantages over GPU deployments; these are company-reported figures and not independently benchmarked here VERIFIED F11. The point is not the vendor. The point is that a serving architecture which halves your megawatts changes which queue you need to be in, and that is now an argument you can make to a permit officer.
Part 3
Evidence Pack B: Proof of Behavior
For any organization running agents inside a supervised process. Eight artifacts. The FDA's 26 questions cluster around exactly this territory, and the same structure answers a bank examiner or an internal audit request VERIFIED F02.
B1. Output sampling protocol
A fixed cadence, a defined sample size and a named reviewer. Weekly is usually right to start. The protocol matters more than the volume; an unsampled system produces no evidence at all.
B2. Version binding
Every sampled output stored against the model version, the prompt scaffold, the tool set and the retrieval index in force at the time. Without this, a drift finding cannot be attributed to a cause.
B3. Drift threshold and trigger
A pre-declared numeric threshold that triggers human review when crossed. Declared in advance is the whole point. A threshold set after an incident is a rationalization.
B4. Decision log
What the agent decided, on what evidence, at what time, under which authority. Queryable, not reconstructable from conversation.
B5. Reversal procedure with a clock
Who can reverse, through which mechanism, and how long it takes measured rather than estimated. The clock is the part that fails under test.
B6. Autonomy baseline and change watch
A recorded baseline of the autonomy setting for every agent tool, plus a subscription to vendor release notes. Claude Code sessions defaulted to auto mode for Pro, Max and Team users from 14 August 2026 unless pinned CITED F15. Vendor defaults change without a purchase order.
B7. Data-classification decision per tier
A written answer for each commercial tier you use about what data may traverse it. Meta's contributor tier trades roughly an order of magnitude of price against training rights over prompts and completions CITED F10. That is a classification decision, not a procurement one.
B8. Infrastructure exposure check
An inventory of AI infrastructure against known-exploited vulnerabilities. CISA added a Ray remote-code-execution flaw, CVE-2025-62593, to the KEV catalog on 17 August 2026 with a three-day federal remediation clock VERIFIED F16. Shadow clusters are the common failure here.
Why completion rate is not evidence
A Princeton-authored paper accepted to ICML 2026 evaluated 15 models across two benchmarks and found capability gains produced only small reliability improvements, proposing 12 metrics across consistency, robustness, predictability and safety. Wider analysis of 15 agent benchmarks concludes that evaluation methodology, not model capability, is the bottleneck on reliable deployment CITED F19. If your program reports task completion and cost per task, you are reporting the metric the research community has just described as insufficient, to an audience that is about to ask for a different one.
Part 4
The shadow-mode graduation test
Idaho National Laboratory's recommended pattern for AI decision support in critical energy environments is phased: run in shadow mode alongside human decision-making before increasing autonomy CITED F22. It generalises cleanly to every regulated sector we work in, and it converts the autonomy debate from a philosophical argument into a measurement.
| Stage | Agent authority | Exit condition | Artifact produced |
|---|---|---|---|
| S1 Shadow | Observes and recommends. No action. Human never sees the recommendation before deciding. | Agreement rate between agent recommendation and human decision stable at a pre-set level across a defined period. | Paired recommendation and decision log. This is your baseline forever. |
| S2 Advisory | Recommendation is visible to the human before the decision. | No increase in decision error rate versus the S1 baseline, plus measured time saved. | Automation-bias check. Watch for the human agreeing more simply because the agent spoke first. |
| S3 Bounded action | Acts inside a defined envelope: value limits, entity limits, reversibility limits. | Reversal procedure tested under load with a measured clock, not an estimated one. | Reversal drill record with elapsed times. |
| S4 Supervised autonomy | Acts by default. Human oversees by exception. | Not an exit, a standing condition. Sampling and drift thresholds from Pack B remain live indefinitely. | Continuous postmarket evidence file. |
The one metric to add this quarter
Recommendation-versus-decision agreement rate, per workflow, trended weekly. It is the exit condition for S1, it is the automation-bias detector at S2, and it is the only number we know of that a regulator, an operator and a finance director all read the same way. Deloitte's finding that 61 percent of leaders expect most agents to be generally autonomous with humans in oversight makes this a near-term requirement rather than a nicety CITED F17.
Part 5
Sector control sheets
Financial services
The situation. Revised US interagency model risk guidance issued 17 April 2026 replaced SR 11-7 after fifteen years and states that generative and agentic AI models are not within its scope, with a request for information signalled to follow VERIFIED F13. Meanwhile the Monetary Authority of Singapore confirmed on 5 August 2026 that autonomous agents sit inside its supervisory expectations CITED F24.
The control. Run agentic use cases through your existing model risk lifecycle voluntarily and document that you chose to. Keep validation artifacts to the SR standard even where scope says they are not required. When the RFI closes and scope arrives, the institution with a two-year artifact trail is having a fundamentally different conversation.
The trap. Treating exclusion as permission. A voice agent taking an approved action on a customer account is doing something a model risk framework would recognize, inside a scope that currently says it does not apply. That asymmetry will not survive the first supervisory finding.
Healthcare
The situation. The FDA discussion paper of 18 August 2026, docket FDA-2026-N-7874, poses 26 questions and proposes a two-axis risk framework, with comments open to 19 October 2026 and an explicit statement that no policy change is proposed VERIFIED F02. There is no comprehensive federal health AI statute; oversight is a patchwork of FDA, CMS and HHS layered with state law.
The control. Stand up Evidence Pack B on your highest-volume generative workflow now, and use the resulting data as the substance of your comment on the docket. An organization that submits real postmarket monitoring data is shaping the framework it will be measured against.
The trap. Assuming administrative equals out of scope. An intake agent that summarizes symptoms for a triage queue has begun influencing clinical judgment, which is the line the framework is drawn around.
Manufacturing
The situation. Humanoid deployment has crossed into measured shift work. Figure AI reports 40 Figure 03 units at BMW Spartanburg after an eleven-month trial with trade press citing ten-hour shifts and more than 90,000 parts handled, and Agility reports Digit past 65,000 hours across nine facilities CITED F20 CITED F21. The reliability figures, including reported placement accuracy above 99 percent, are company and trade-press reported and not independently audited CITED F20.
The control. Write the acceptance test before the purchase order. Specify the cell, part mix, shift length and failure taxonomy, then measure a full production quarter against your own baseline. Vendor operating hours are a marketing metric; your hours on your parts are an operating metric.
The trap. Underwriting an uptime commitment to a customer against a number nobody outside the vendor has verified.
Energy
The situation. Load is the limiter, not compute. Texas is auditing roughly 250 to 300 projects against more than 474 GW requested VERIFIED F04, Pennsylvania made grid commitments a permit condition VERIFIED F03, and roughly 28 percent of the 1,066 GW requested nationally is expected to convert CITED F05. Operationally, Argonne's GridMind and the Idaho National Laboratory phased pattern point toward governed co-pilots rather than autonomous control CITED F22.
The control. Shadow mode as the mandatory first stage for every operational agent, with a written graduation test. Air-gapped deployment remains the pattern for nuclear and critical control environments.
The trap. Letting a vendor demonstration set the autonomy level. The demonstration is S4. Your first deployment is S1.
Part 6
The fourteen-point score sheet
Score each item 0 for absent, 1 for partial and documented, 2 for complete and dated. Run it in one hour with the people who actually operate the system, not the people who sponsored it. Our working floor for a regulated deployment is 20 out of 28, with no zero on items 9 through 12.
| # | Item | Evidence that scores 2 |
|---|---|---|
| 1 | Interconnection queue position is documented | Queue name, position, batch ID, dated study schedule |
| 2 | Firm generation is contracted or under build | Signed agreement with capacity and commissioning date |
| 3 | Infrastructure cost allocation is written | Named party and amount for network upgrades |
| 4 | Local approval status is evidenced | Dated approval plus objections and responses |
| 5 | Audit or pause exposure is stated in writing | Explicit yes or no from the developer, dated |
| 6 | Capacity conversion assumption is explicit | Stated probability with the reasoning attached |
| 7 | Output sampling runs on a fixed cadence | Protocol document, named reviewer, last four sample sets |
| 8 | Version binding is complete | Model, prompt scaffold, tools and index stored per sample |
| 9 | Drift threshold is declared in advance | Numeric threshold with a dated approval before go-live |
| 10 | Decision log is queryable | A query returns decision, evidence, time and authority |
| 11 | Reversal clock is measured, not estimated | Drill record with elapsed times from the last quarter |
| 12 | Autonomy baseline is recorded and watched | Per-tool setting baseline plus release-note subscription |
| 13 | Data classification is decided per commercial tier | Written decision per tier, signed by a data owner |
| 14 | AI infrastructure is checked against KEV | Dated inventory reconciled to the CISA catalog |
How to read your score
- 24 to 28. You can answer a supervisory request from a query rather than a project. Keep the sampling cadence live and re-score quarterly.
- 20 to 23. Workable. Find the zeros first; a single zero on items 9 to 12 is worth more attention than three ones elsewhere.
- 14 to 19. You have a governance document and not a governance system. Start with item 10, the decision log, because most of the others become cheap once it exists.
- Below 14. Do not widen autonomy on any workflow until items 9 to 12 are non-zero. Deloitte found only 5 percent of organizations describe their processes as highly prepared for agents, so this band is crowded and not disqualifying CITED F17.
Part 7
Questions, answers, and four things to ask your partner
Is any of this legally required today?
Parts of it. EU AI Act Article 50 transparency obligations have been applicable since 2 August 2026 with penalties reaching EUR 15 million or 3 percent of worldwide annual turnover, and a transitional period to 2 December 2026 for marking and detection on generative systems already on the market. Annex III high-risk obligations are placed at December 2027 and Annex I at August 2028 under the Digital Omnibus; commentary asserting high-risk applicability from 2 August 2026 is contested and we follow the primary legal sources FLAG CITED F14. Pennsylvania's permit conditions are binding now VERIFIED F03. The FDA paper changes no policy VERIFIED F02. Most of Pack B is not yet required anywhere, which is exactly why building it now is cheap.
Should we pause agent work until the rules settle?
The evidence points the other way. Financial services has an explicit exclusion VERIFIED F13 and healthcare has a paper that expressly changes nothing VERIFIED F02. Neither is a moratorium. Both are windows. Organizations that spend the window producing evidence will describe a working system when rules arrive. Organizations that pause will describe an intention.
Our vendor says they handle governance. Is that enough?
Vendor controls are necessary and not sufficient. Capital is flowing into agent visibility precisely because buyers are stuck here; Obsidian Security raised at a USD 1.1 billion valuation in early August on that thesis CITED F25. Useful tooling. But a regulator will ask you what your agent decided and who could reverse it, and a vendor dashboard you cannot export is not an answer.
What does this cost to stand up?
Pack A is mostly document requests to parties you already pay, so the cost is coordination rather than capital. Pack B items 1 to 5 are typically a small engineering effort against systems you already have. The expensive item is the decision log where none exists, and it is expensive once. We have not seen a case where the full pack cost more than the first supervisory finding it would have prevented.
Four questions for your data-center partner, before signing
- What is our queue position, in which batch, and what is the current study date in writing?
- What firm generation serves this load, under what contract, commissioning when?
- Who funds the network upgrades, and is that allocation written into our agreement?
- Is this project inside any current audit or pause cohort, yes or no, dated?
If any answer arrives as a reassurance rather than a document, treat it as a zero on the score sheet.
Where this fits
These two packs are the evidence layer of the AEGIS Framework, the Agentic Enterprise Governance and Intelligence Standard. Pack A and Pack B are what an AEGIS Diagnostic produces in its first fortnight, and what AEGIS Run keeps alive afterwards. You do not need the framework to use the packs. The packs work on their own, which is the point of publishing them.
AI governance at Ariana.Digital · AI Readiness Brief · myndQ
Sources
Field note claim groups use the F prefix and map one to one onto the C-numbered research base published with the Daily Market Pulse edition of 2026-08-21.
- F02 FDA, Considerations for the Regulation of Generative AI-Enabled Medical Devices, 18 August 2026, docket FDA-2026-N-7874, 26 questions, two-axis risk framework, comments to 19 October 2026, no policy change proposed. FDA · FDA press announcement
- F03 Pennsylvania Executive Order 2026-05, signed 18 August 2026: binding GRID commitment and local approval as permit preconditions, developer-funded infrastructure, removal from Fast Track permitting, prohibition on non-disclosure agreements. Executive Order text · Commonwealth release
- F04 Texas Governor Abbott directive of 3 August 2026 pausing pending data-center approvals for audit; more than 474 GW proposed load, roughly 250 to 300 projects audited; ERCOT sought a good-cause exception on the 7 August Batch Zero deadline. Utility Dive · Holland & Knight
- F05 Wood Mackenzie analysis reported by Bloomberg, 12 August 2026: roughly 28 percent of 1,066 GW requested expected to be committed, leaving about 768 GW speculative or duplicative. Bloomberg
- F10 Meta Muse Code, 5 August 2026, contributor tier at approximately USD 0.10 and USD 0.20 per million tokens in exchange for training rights over prompts and completions against standard USD 1.25 and USD 4.25. CNBC
- F11 Cerebras CS-4 announced 18 August 2026; company claims up to twice CS-3 speed, up to 30 times GPU tokens-per-second-per-user and up to 10 times CS-3 throughput per watt. Company-reported. Cerebras investor release
- F13 Federal Reserve, OCC and FDIC revised model risk management guidance, 17 April 2026, replacing SR 11-7 and stating generative and agentic AI models are not within scope; request for information planned. OCC · American Banker
- F14 EU AI Act Article 50 applicable from 2 August 2026, penalties to EUR 15 million or 3 percent of worldwide annual turnover, transition to 2 December 2026 for marking and detection; Annex III high-risk placed at December 2027 and Annex I at August 2028 under the Digital Omnibus. Contrary market commentary is contested. Article 50 · European Commission FAQ
- F15 Claude Code auto mode default for Pro, Max and Team from 14 August 2026 unless pinned, replacing repeated approval prompts with a classifier screening tool calls. 9to5Mac · Anthropic newsroom
- F16 CISA added CVE-2025-62593, a Ray remote-code-execution flaw affecting versions prior to 2.52.0 with CVSS 4.0 score 9.4, to the Known Exploited Vulnerabilities catalog on 17 August 2026 with a federal remediation deadline of 20 August 2026. CISA KEV catalog · The Hacker News
- F17 Deloitte agentic AI research, August 2026, 501 US respondents across five industries fielded April to June 2026: 74 percent expect nearly half of processes redesigned within four years, 61 percent expect most agents generally autonomous with human oversight, 5 percent report processes highly prepared. Deloitte Insights
- F19 Princeton-authored ICML 2026 paper evaluating 15 models across two benchmarks with 12 proposed reliability metrics; wider analysis of 15 agent benchmarks identifying evaluation methodology as the primary bottleneck. Standardized AI evaluation · Springer review
- F20 Figure AI: 40 Figure 03 units at BMW Spartanburg after an eleven-month trial, trade press reporting ten-hour shifts, more than 90,000 parts and above 99 percent placement accuracy. Company and trade-press reported, not independently audited. Technology.org · Humanoid Guide
- F21 Agility Robotics reports Digit past 65,000 operating hours across nine customer facilities. Company-reported. Solid Market Research
- F22 Argonne National Laboratory GridMind operator co-pilot; Idaho National Laboratory phased deployment running decision support in shadow mode before increasing autonomy; air-gapped pattern for nuclear and critical control environments. Grid co-pilot analysis · POWER Magazine
- F24 Monetary Authority of Singapore written parliamentary reply on agentic AI in financial services, 5 August 2026, confirming autonomous agents fall inside supervisory expectations. MAS
- F25 Obsidian Security USD 85 million Series D at a USD 1.1 billion valuation, 4 August 2026, extending agent governance controls to Claude Code and Claude Cowork. SecurityWeek