Quality audit · Verdal Group (fictional case)

L3 — Dual external audit

Two models outside the system — Claude Fable 5 and GPT-5.6 Sol — given the identical prompt, neither seeing the other.

QA Level 3 — Dual external audit, WP1 Market

Documents audited: livrables/verdal-wp1-market.pptx (17 slides) + livrables/verdal-wp1-market-model.xlsx (7 tabs) Date: 27 July 2026 · Engagement: VG-STRAT-2026-07 · Stakes: CRITICAL (Board capital allocation) Models: Claude Fable 5 (claude-fable-5, Anthropic API) and GPT-5.6 Sol (gpt-5.6-sol, OpenAI API) — the consultant's explicit choice, both confirmed present on their respective APIs before the run. Both received the identical prompt, context and full document package; neither saw the other's output, nor this session's L1 or L2 findings. Package sent: full extracted text of all 17 slides plus every model cell, formula and computed value (~52,000 characters).

No deliverable was modified.

Run transparency


Scores

# Criterion Fable 5 GPT-5.6 Sol Δ Convergence
1 Internal coherence (deck↔︎model, deck↔︎itself) 3 1 2 ⚠️ divergence — HIGH
2 Conformity to the firm's standards 3 2 1 ✓ converge (weak)
3 Voice and tone consistency 4 3 1 ✓ converge
4 Source quality and traceability 4 2 2 ⚠️ divergence — HIGH
5 Analytical rigour of the sizing 4 1 3 ⚠️ deep divergence — HIGH
6 Rigour of the €25bn test (folded into #5, 4) 1 3 ⚠️ deep divergence — HIGH
7 Treatment of uncertainty 4 3 1 ✓ converge
8 Actionability for a Board 4 2 2 ⚠️ divergence — MEDIUM
9 Arithmetic integrity 4 3 1 ✓ converge
10 Model architecture, transparency, control not scored 3 GPT-only criterion
11 Economic / financial modelling rigour (NPV, IRR, cash) not scored 1 GPT-only criterion
12 Conceptual integrity of the value-chain perimeter not scored 2 GPT-only criterion
D1 Colour / charter coherence 3 (tbc render) 3 (tbc render) 0 ✓ converge
D2 Typographic hierarchy and legibility 3 (tbc render) 3 (tbc render) 0 ✓ converge
D3 Content → visual fit 4 3 1 ✓ converge
D4 Density and breathing room 4 2 2 ⚠️ divergence — MEDIUM
D5 Quality of structured visuals 3 3 0 ✓ converge
D6 Template / master consistency 3 (tbc render) 2 1 ✓ converge
D7 Exec legibility and narrative 5 2 3 ⚠️ deepest divergence — HIGH
D8 Finish 3 2 1 ✓ converge

Combined score. Fable 5: 3.5/5 (self-reported; 58/80 computed across its 16 criteria). GPT-5.6 Sol: 2.2/5 (44/100 across its 20 criteria; no self-reported overall).

Verdicts are opposed and that is the finding.

Fable 5: "a genuinely strong WP1 package with a weak seam between its two artefacts… None of these blocks delivery, all of them are fixable in a day… Fix the deck to the model's standard of honesty and this clears the partner plausibility test."

GPT-5.6 Sol: "HOLD. The package is not fit for interim Board delivery in its current form. The central mandate question is presented as answered when the source of the €25bn figure is still unknown… the analysis is a highly assumption-led estate illustration rather than a genuine addressable-market or capital-allocation model. The candour in the workbook is materially better than the certainty of the deck."

They agree on almost every specific defect. They disagree on whether the specific defects are the whole problem. That is a scope judgement, not a scoring artefact, and it is the consultant's to make.


Convergences — both models agree

Confirmed strengths

  1. The workbook's honesty layer. Both single it out: three-tier confidence on all 30 inputs, the yellow-fill convention, A44/A45 naming the two cells that flip the answer, F37 disclosing that the Savills evidence is internally contested, and the explicit statement that no customer research exists. GPT: "explicitly identifies material unsourced assumptions." Fable: "near best-practice."
  2. The perimeter architecture is the right answer to the mandate. Both accept the four-level taxonomy and that only level 4 is Verdal's money. GPT: "correctly distinguishes broad European charging expenditure from the much smaller economics potentially retained by a property host."
  3. The core arithmetic ties. Both independently reproduced the funnel, the volume build, the three margin stacks and the €70m cost stack. Fable recomputed the chain end to end and recorded it "for the record".
  4. No invented Verdal internal figure. Both note that unsupplied parameters are routed to data-room requests rather than estimated.
  5. The footfall counterargument is deliberately surfaced, not buried. Fable calls it "the strategically honest move" — the package flags that its own P&L may be the wrong test and routes it to WP5.

Confirmed weaknesses — fix these first

# Finding Fable 5 GPT-5.6 Sol Also found at L2
C1 Slide 13's "Charger spec" and "Bays per charger" bars exist in no model cell. Assumptions!F17 itself admits the sensitivity "belongs in the sensitivity and it is not in it." "Two bars on a Board exhibit are computed off-model — indefensible if the Board asks to see the cells." "Slide 13 displays sensitivities that are not exposed as dedicated outputs on tab 5, preventing direct deck-to-cell traceability."
C2 FX error on the Savills conversion. £3,000-5,000 × 4,400 bays = **£**13.2-22.0m, presented as "€13-22m". At £1 = €1.17 it is €15.4-25.7m; slide 11's "2-3x" is 2.3-3.9x. ✓ explicit ✓ explicit
C3 Slide 8's step percentages are computed from rounded display labels, not from tab 1. Shown −69.4% and −42.3%; true −70.0% and −43.0%. ✓ explicit (implicit, via precision critique)
C4 The confidence scale is applied rigorously in the model and appears on no slide — precisely the client's stated scoring criterion.
C5 Slide 2's "on a destination-grade 75 kW unit it is reached today" is an overclaim. The model shows about break-even at 2030 modelled throughput on an unsourced 80 kWh/bay/day Hypothesis. "an overclaim of exactly the kind standard 11 forbids" "not evidence of current utilisation"
C6 The lease-fee flip condition has no visual anywhere in the deck. A44 says at the Savills range "the recommendation flips"; slide 11's headline still asserts the revenue share wins. ✓ listed as a critical issue ✓ listed as a risk
C7 The break-even is not a whole-estate solve. tab 5!B9 = 159.8 kWh/bay/day is DC-only; the full P&L is still −€0.6m at 160. (not raised) "full break-even is approximately 162 kWh per DC bay per day"
C8 The 27 July build date precedes the 4 August kick-off. Both raise it as a document-control and credibility issue for a Board-facing artefact.
C9 The pilot is asked for without a budget, scope or go/no-go threshold, so the 8 September "pilot / no pilot" decision is not yet a real decision.

Divergences — ranked by criticality

D-1 · Analytical rigour of the sizing — Fable 4, GPT 1 · HIGH · GPT is better supported

GPT's case. "It assumes all 1,100 stores can support four bays and treats this theoretical estate footprint as addressable volume. It does not screen site ownership or lease restrictions, car-park size, grid headroom, planning constraints, operator exclusivity, country economics, rollout timing or economically viable site thresholds." GPT wants a TAM/SAM/SOM structure with a site-archetype table replacing the single four-bay placeholder.

Fable's case. The architecture is sound — top-down funnel bounded by an independent bottom-up build with an explicit reconciliation corridor, and the bay count is correctly flagged as the single replaceable cell and data-room request no. 1.

Adjudication: GPT. Fable is judging the model's construction; GPT is judging whether "genuinely addressable" — the mandate's own word — has been established. It has not. And GPT's objection is independently corroborated by the L1 source pass: the Savills benchmark that would flip the recommendation is quoted only for sites with "proximity to major road networks / 15,000+ vehicles passing by daily / nearby amenities". The share of 1,100 grocery car parks that clears that bar is unknown, unmodelled, and it bounds both the lease upside and the operate case. A site-eligibility screen is a real gap, not a stylistic preference.

Note on proportionality: WP3 owns country prioritisation and WP5 owns entry mode. A full site screen is not WP1's job. But WP1 currently calls 1,100 stores × 4 bays "addressable" without saying that eligibility has not been tested — and that sentence is one line, not a work package.

D-2 · Exec legibility and narrative — Fable 5, GPT 2 · HIGH · both are right about different things

Fable: "The title-flow test passes cleanly… This is the deck's best dimension." GPT: "the title flow does not form a reliable argument because the core €25bn conclusion is unproven and several titles overstate conditional analysis."

Adjudication: split, and it resolves. Fable and the L2 subagent both ran the title-flow test independently and both passed it — mechanically, the 12 titles do carry the whole argument. GPT is not scoring the mechanics; it is scoring whether a title flow built on an unproven premise can be called reliable. That objection is the coherence finding (D-4 below) counted twice. Score the mechanics as Fable does; act on the premise as GPT does. The correct reading is: the narrative instrument is excellent and it is currently carrying one claim it cannot support.

GPT adds one point Fable does not: the deck ends with activity requests, not a Board decision. "End with a governing recommendation such as: 'Do not allocate rollout capital at the interim; authorise a capped pilot and lease-term renegotiation once five named evidence gaps are closed.'" That is a real improvement and it is free.

D-3 · Source quality and traceability — Fable 4, GPT 2 · HIGH · GPT is better supported, and L1 settles it

Fable praises the Assumptions tab's discipline; GPT says the deck's source lines are not traceable enough for a critical Board document — no report dates, no page references, no URLs, no quoted market definitions, no currency or horizon basis on slide 5.

Adjudication: GPT, decisively — the L1 pass found six contradicted claims and five confidence labels needing downgrade. ACEA is not "of the same order" as Société Générale (€8bn annual vs €30bn cumulative); the Grand View 2030 figure is no longer reproducible at source; Market Data Forecast's US25.4bnbelongstoits * charger * report, whileits * chargingstation * reportsaysUS671bn for the same year; the ADAC primary could not be reached at all. Fable did catch the currency and horizon-year mismatch on slide 5 and the FX error — so it saw the symptoms and under-weighted them.

D-4 · Internal coherence — Fable 3, GPT 1 · HIGH · GPT is better supported

GPT: "Slides 4 and 6 and model tab 4 cell C5 treat the Board's €25bn as cumulative infrastructure build-out, while… model tab 1 cell A20 says annual plug revenue is the most likely interpretation."

Adjudication: GPT, and the L2 audit found the identical contradiction independently. Fable scored coherence on four other contradictions and did not surface this one. On the single question the client commissioned, the package holds two answers. This is the highest-value finding in the whole L3 pass, and it was found twice out of three independent auditors.

GPT's proposed fix is better than a choice: "Until the Board-paper citation is obtained, rewrite slides 4-6 as scenarios: 'If the source is build-out capex…', 'If it is annual plug revenue…', 'If it is the McKinsey mobility-finance report…'" That is more honest than picking one, and it is exactly what the client said it would score on.

D-5 · Actionability for a Board — Fable 4, GPT 2 · MEDIUM · partial convergence hidden by the gap

GPT wants an investment case: NPV, IRR, cash requirement, rollout profile, hurdle rate, staged commitment. Fable calls slides 15-16 "concretely actionable" and the "threshold, not forecast" framing "exactly right for a capital-allocation Board at this stage".

Adjudication: Fable on scope, GPT on the gap. The WP1 methodology asks for market sizing, three margin stacks and a bounded range — not for an investment case. NPV and IRR belong to WP5, which scores entry mode. GPT is applying a criterion the mandate does not set for this work package. But GPT is right that a Board asked to allocate capital will want it, and right that the pilot ask has no number attached. Both models independently say: scope and cost the pilot.

D-6 · Density — Fable 4, GPT 2 · MEDIUM · converge on the specifics

Both name slides 5 and 16 as prose-heavy. They differ on the separators: GPT calls slides 3, 7, 10 and 17 "sparse" and slide 17 "effectively an empty placeholder"; Fable treats the separators as template furniture. Measured: slide 17 uses the Fin layout, whose master provides Nom / Rôle / Email placeholders that the slide does not instantiate — so nothing renders empty, but the closing page carries only "WP1 — Market" with no doc ID, confidentiality line or leave-behind. Fable's fix is the better one: "finish slide 17 with the doc ID, confidentiality line and the three data-room requests as a leave-behind recap."

D-7 · Source line on the separator pages — LOW · a genuine reading of the firm's own rule

GPT scores a standards breach: "Slides 3, 7, 10 and 17 have no source line despite the no-exception rule." Fable and the L2 audit both treat dividers and book-ends as correctly exempt. The firm's standard says "Every page carries its source line. No exception." This is a consultant call on the firm's own rule, not a defect, and it should be settled once at cabinet level rather than re-argued each deck.


Points each model found alone — verified against the files

Fable 5 only, and correct

Finding Verification
Slides 15 and 16 claim the lease fee is "the only" Verdal figure not supplied — and both pages contradict themselves. Slide 15 item 2: "the only number in this analysis that Verdal has and we do not"; slide 16: "the only Verdal figure we had to invent". Confirmed. The same slide 15 lists item 1 (convertible bays, "Held by Verdal today") and item 4 (energy contract, "Held by Verdal today") as also held and unsupplied. Three, not one. A factual self-contradiction on the two most action-oriented pages.
tab 5!B31 contribution = €0.37/kWh (gross of payment costs) while tab 5!B6 = €0.34 (net). Slide 13's "Contribution/kWh €0.25 to €0.45" does not say which definition it varies. Confirmed. Two contribution definitions live on the same tab.
Slide 6's displayed steps do not tie to each other: 20,356 − 19,590 = 766; 766 − 718 = 48, against €47.6m in the model. Rounding applied per bar rather than from unrounded cells. Confirmed. Same defect class as slide 8's percentages.
Slide 12's three largest cost lines (€48.2m of €70m) rest on a single adviser source — Interpath, May 2026 — with no independent second source. Confirmed, and L1 adds that Interpath's figures are its own chart-footnote modelling assumptions, not operator disclosures. Interpath's own footnote assumes €180,000 per charger; the model uses €200,000.
The 5.8% reconciliation is partly circular — both routes share the same unsourced throughput assumptions, so "agreement" is weaker evidence than it looks. Fair. tab 2!D17 says as much in the model; the slide does not.

GPT-5.6 Sol only, and materially important

Finding Verification
Possible double count between the Eurostat price and the grid capacity adder. Eurostat's non-household band IC price includes network charges; the model adds €0.04/kWh of "grid capacity and demand charges" on top. Real risk, worth €3.37m a year on slide 12. Not a certain error — a large DC hub's medium-voltage capacity charge is billed on kW and is genuinely additional — but the model nowhere separates commodity, volumetric network, and fixed capacity charges. This needs one line of disclosure at minimum, and probably a decomposition.
Band mismatch worth checking. 84.315 GWh across 1,100 sites = 76.6 MWh per site per year, which is below the 500-2,000 MWh floor of Eurostat band IC. Measured and confirmed. It is defensible if charging shares the store's existing supply point (a supermarket's total load sits comfortably in band IC or above) — but that is an assumption the model does not state. One sentence in F24.
The €20.4bn plug-revenue figure applies an unweighted (DC + AC)/2 = €0.525/kWh to all European public charging, when public charging is DC-dominated by energy. Confirmed. At a 70/30 DC/AC energy mix the figure is ~€21.5bn. Small in absolute terms; it undercuts the "precision" the number is presented with.
No page numbers on any slide. Measured and confirmed — GPT is right and Fable is wrong here. The template carries no footer or slide-number placeholder on any of its four layouts. Slides 2-17 have no page number at all. (Fable read the "1" on slide 2 as a page number; it is the first support badge.) For a Board deck this is a real navigability defect — nobody can say "turn to page 12".
Verdal is not only a site host under self-operation — it also becomes the charge-point operator, so slide 4's "Verdal, today and under every structure tested" on the level-4 row is wrong. Confirmed and conceptually clean. Under self-operation Verdal occupies levels 3 and 4.
Describing most plug revenue as energy-supplier pass-through is imprecise — the CPO retains the spread and pays only its delivered energy cost. Fair. tab 1!F14 and slide 4 both use the pass-through framing; it is a simplification a Board member in energy will notice.

Fable's two "confirm at render" concerns — resolved by measurement


Priority recommendations from L3

Tier 1 — before the deck is shown to anyone (both models, plus L2)

  1. Resolve the €25bn perimeter split between deck and model, or convert slides 4-6 to explicit conditional scenarios per GPT's proposal.
  2. Either build the charger-spec and bays-per-charger sensitivities as live cells in tab 5, or strip those two bars from slide 13. One or the other must happen.
  3. Correct the FX on the Savills range (F37 and slide 11) and state the rate used.
  4. Recompute slide 6's and slide 8's displayed steps from unrounded cells.
  5. Re-label slide 6's third bridge step — it is the operator's retained 85%, not a cost of delivery.
  6. Delete "only" from slides 15 and 16.

Tier 2 — before the interim on 8 September 7. Put the confidence taxonomy on the slides. 8. Give the lease-fee flip condition a visual on slide 11, not a bullet. 9. Solve the break-even against the whole estate (162.2, not 159.8) or re-label B9 as DC-only. 10. Disclose or decompose the possible network-charge double count, and state the Eurostat band basis. 11. State that site eligibility has not been screened — one line on slide 9, and a WP3 hand-off. 12. Scope and cost the pilot so 8 September is a real budget decision. 13. Re-date and version the deck for the interim. 14. Add page numbers.

Tier 3 — improvements, not defects 15. Slide 5 as a table rather than three prose blocks, with currencies and horizon years visible. 16. Slide 2 from six supports to three or four. 17. End on a governing recommendation rather than a list of requests. 18. Correct slide 4's level-4 row: under self-operation Verdal is both operator and host.


L3 verdict

Both models scored the same package and reached opposite delivery verdicts — 3.5/5 and "fixable in a day" against 2.2/5 and "HOLD". That divergence is not noise. It is the difference between judging the package against the WP1 methodology (Fable's frame, where the mandate asks for a sizing, three margin stacks and a bounded range, and gets them) and judging it against what a Board needs to allocate capital (GPT's frame, where the absence of a site screen, an investment case and a decision ask is disqualifying).

Both frames are legitimate and the choice between them is the consultant's. The honest reading is that Fable is right about WP1 and GPT is right about the engagement: the work package has done what it was scoped to do, and the deck currently presents it with more certainty than the model supports.

What is not in dispute: nine specific defects were found by both models independently, six of them also by the internal L2 pass. Those six should be fixed regardless of which frame the consultant adopts, and none of them requires new research.