Two models outside the system — Claude Fable 5 and GPT-5.6 Sol — given the identical prompt, neither seeing the other.
Documents audited:
livrables/verdal-wp1-market.pptx (17 slides) +
livrables/verdal-wp1-market-model.xlsx (7 tabs)
Date: 27 July 2026 · Engagement:
VG-STRAT-2026-07 · Stakes: CRITICAL (Board capital allocation)
Models: Claude Fable 5
(claude-fable-5, Anthropic API) and GPT-5.6
Sol (gpt-5.6-sol, OpenAI API) — the consultant's
explicit choice, both confirmed present on their respective APIs before
the run. Both received the identical prompt, context
and full document package; neither saw the other's output, nor this
session's L1 or L2 findings. Package sent: full
extracted text of all 17 slides plus every model cell, formula and
computed value (~52,000 characters).
No deliverable was modified.
claude-fable-5 rejects the
temperature parameter
(temperature is deprecated for this model); the run was
re-issued without it. GPT-5.6 Sol ran with default sampling and
max_output_tokens: 32000.background: true mode.| # | Criterion | Fable 5 | GPT-5.6 Sol | Δ | Convergence |
|---|---|---|---|---|---|
| 1 | Internal coherence (deck↔︎model, deck↔︎itself) | 3 | 1 | 2 | ⚠️ divergence — HIGH |
| 2 | Conformity to the firm's standards | 3 | 2 | 1 | ✓ converge (weak) |
| 3 | Voice and tone consistency | 4 | 3 | 1 | ✓ converge |
| 4 | Source quality and traceability | 4 | 2 | 2 | ⚠️ divergence — HIGH |
| 5 | Analytical rigour of the sizing | 4 | 1 | 3 | ⚠️ deep divergence — HIGH |
| 6 | Rigour of the €25bn test | (folded into #5, 4) | 1 | 3 | ⚠️ deep divergence — HIGH |
| 7 | Treatment of uncertainty | 4 | 3 | 1 | ✓ converge |
| 8 | Actionability for a Board | 4 | 2 | 2 | ⚠️ divergence — MEDIUM |
| 9 | Arithmetic integrity | 4 | 3 | 1 | ✓ converge |
| 10 | Model architecture, transparency, control | not scored | 3 | — | GPT-only criterion |
| 11 | Economic / financial modelling rigour (NPV, IRR, cash) | not scored | 1 | — | GPT-only criterion |
| 12 | Conceptual integrity of the value-chain perimeter | not scored | 2 | — | GPT-only criterion |
| D1 | Colour / charter coherence | 3 (tbc render) | 3 (tbc render) | 0 | ✓ converge |
| D2 | Typographic hierarchy and legibility | 3 (tbc render) | 3 (tbc render) | 0 | ✓ converge |
| D3 | Content → visual fit | 4 | 3 | 1 | ✓ converge |
| D4 | Density and breathing room | 4 | 2 | 2 | ⚠️ divergence — MEDIUM |
| D5 | Quality of structured visuals | 3 | 3 | 0 | ✓ converge |
| D6 | Template / master consistency | 3 (tbc render) | 2 | 1 | ✓ converge |
| D7 | Exec legibility and narrative | 5 | 2 | 3 | ⚠️ deepest divergence — HIGH |
| D8 | Finish | 3 | 2 | 1 | ✓ converge |
Combined score. Fable 5: 3.5/5 (self-reported; 58/80 computed across its 16 criteria). GPT-5.6 Sol: 2.2/5 (44/100 across its 20 criteria; no self-reported overall).
Verdicts are opposed and that is the finding.
Fable 5: "a genuinely strong WP1 package with a weak seam between its two artefacts… None of these blocks delivery, all of them are fixable in a day… Fix the deck to the model's standard of honesty and this clears the partner plausibility test."
GPT-5.6 Sol: "HOLD. The package is not fit for interim Board delivery in its current form. The central mandate question is presented as answered when the source of the €25bn figure is still unknown… the analysis is a highly assumption-led estate illustration rather than a genuine addressable-market or capital-allocation model. The candour in the workbook is materially better than the certainty of the deck."
They agree on almost every specific defect. They disagree on whether the specific defects are the whole problem. That is a scope judgement, not a scoring artefact, and it is the consultant's to make.
A44/A45 naming the two cells that flip the
answer, F37 disclosing that the Savills evidence is
internally contested, and the explicit statement that no customer
research exists. GPT: "explicitly identifies material unsourced
assumptions." Fable: "near best-practice."| # | Finding | Fable 5 | GPT-5.6 Sol | Also found at L2 |
|---|---|---|---|---|
| C1 | Slide 13's "Charger spec" and "Bays per charger" bars exist
in no model cell. Assumptions!F17 itself admits
the sensitivity "belongs in the sensitivity and it is not in
it." |
"Two bars on a Board exhibit are computed off-model — indefensible if the Board asks to see the cells." | "Slide 13 displays sensitivities that are not exposed as dedicated outputs on tab 5, preventing direct deck-to-cell traceability." | ✅ |
| C2 | FX error on the Savills conversion. £3,000-5,000 × 4,400 bays = **£**13.2-22.0m, presented as "€13-22m". At £1 = €1.17 it is €15.4-25.7m; slide 11's "2-3x" is 2.3-3.9x. | ✓ explicit | ✓ explicit | ✅ |
| C3 | Slide 8's step percentages are computed from rounded display
labels, not from tab 1. Shown −69.4% and −42.3%;
true −70.0% and −43.0%. |
✓ explicit | (implicit, via precision critique) | ✅ |
| C4 | The confidence scale is applied rigorously in the model and appears on no slide — precisely the client's stated scoring criterion. | ✓ | ✓ | ✅ |
| C5 | Slide 2's "on a destination-grade 75 kW unit it is reached today" is an overclaim. The model shows about break-even at 2030 modelled throughput on an unsourced 80 kWh/bay/day Hypothesis. | "an overclaim of exactly the kind standard 11 forbids" | "not evidence of current utilisation" | — |
| C6 | The lease-fee flip condition has no visual anywhere in the
deck. A44 says at the Savills range "the
recommendation flips"; slide 11's headline still asserts the
revenue share wins. |
✓ listed as a critical issue | ✓ listed as a risk | ✅ |
| C7 | The break-even is not a whole-estate solve.
tab 5!B9 = 159.8 kWh/bay/day is DC-only; the full P&L
is still −€0.6m at 160. |
(not raised) | "full break-even is approximately 162 kWh per DC bay per day" | ✅ |
| C8 | The 27 July build date precedes the 4 August kick-off. Both raise it as a document-control and credibility issue for a Board-facing artefact. | ✓ | ✓ | — |
| C9 | The pilot is asked for without a budget, scope or go/no-go threshold, so the 8 September "pilot / no pilot" decision is not yet a real decision. | ✓ | ✓ | — |
GPT's case. "It assumes all 1,100 stores can support four bays and treats this theoretical estate footprint as addressable volume. It does not screen site ownership or lease restrictions, car-park size, grid headroom, planning constraints, operator exclusivity, country economics, rollout timing or economically viable site thresholds." GPT wants a TAM/SAM/SOM structure with a site-archetype table replacing the single four-bay placeholder.
Fable's case. The architecture is sound — top-down funnel bounded by an independent bottom-up build with an explicit reconciliation corridor, and the bay count is correctly flagged as the single replaceable cell and data-room request no. 1.
Adjudication: GPT. Fable is judging the model's construction; GPT is judging whether "genuinely addressable" — the mandate's own word — has been established. It has not. And GPT's objection is independently corroborated by the L1 source pass: the Savills benchmark that would flip the recommendation is quoted only for sites with "proximity to major road networks / 15,000+ vehicles passing by daily / nearby amenities". The share of 1,100 grocery car parks that clears that bar is unknown, unmodelled, and it bounds both the lease upside and the operate case. A site-eligibility screen is a real gap, not a stylistic preference.
Note on proportionality: WP3 owns country prioritisation and WP5 owns entry mode. A full site screen is not WP1's job. But WP1 currently calls 1,100 stores × 4 bays "addressable" without saying that eligibility has not been tested — and that sentence is one line, not a work package.
Fable: "The title-flow test passes cleanly… This is the deck's best dimension." GPT: "the title flow does not form a reliable argument because the core €25bn conclusion is unproven and several titles overstate conditional analysis."
Adjudication: split, and it resolves. Fable and the L2 subagent both ran the title-flow test independently and both passed it — mechanically, the 12 titles do carry the whole argument. GPT is not scoring the mechanics; it is scoring whether a title flow built on an unproven premise can be called reliable. That objection is the coherence finding (D-4 below) counted twice. Score the mechanics as Fable does; act on the premise as GPT does. The correct reading is: the narrative instrument is excellent and it is currently carrying one claim it cannot support.
GPT adds one point Fable does not: the deck ends with activity requests, not a Board decision. "End with a governing recommendation such as: 'Do not allocate rollout capital at the interim; authorise a capped pilot and lease-term renegotiation once five named evidence gaps are closed.'" That is a real improvement and it is free.
Fable praises the Assumptions tab's discipline; GPT says the deck's source lines are not traceable enough for a critical Board document — no report dates, no page references, no URLs, no quoted market definitions, no currency or horizon basis on slide 5.
Adjudication: GPT, decisively — the L1 pass found six contradicted claims and five confidence labels needing downgrade. ACEA is not "of the same order" as Société Générale (€8bn annual vs €30bn cumulative); the Grand View 2030 figure is no longer reproducible at source; Market Data Forecast's US25.4bnbelongstoits * charger * report, whileits * chargingstation * reportsaysUS671bn for the same year; the ADAC primary could not be reached at all. Fable did catch the currency and horizon-year mismatch on slide 5 and the FX error — so it saw the symptoms and under-weighted them.
GPT: "Slides 4 and 6 and model tab 4 cell C5 treat the Board's €25bn as cumulative infrastructure build-out, while… model tab 1 cell A20 says annual plug revenue is the most likely interpretation."
Adjudication: GPT, and the L2 audit found the identical contradiction independently. Fable scored coherence on four other contradictions and did not surface this one. On the single question the client commissioned, the package holds two answers. This is the highest-value finding in the whole L3 pass, and it was found twice out of three independent auditors.
GPT's proposed fix is better than a choice: "Until the Board-paper citation is obtained, rewrite slides 4-6 as scenarios: 'If the source is build-out capex…', 'If it is annual plug revenue…', 'If it is the McKinsey mobility-finance report…'" That is more honest than picking one, and it is exactly what the client said it would score on.
GPT wants an investment case: NPV, IRR, cash requirement, rollout profile, hurdle rate, staged commitment. Fable calls slides 15-16 "concretely actionable" and the "threshold, not forecast" framing "exactly right for a capital-allocation Board at this stage".
Adjudication: Fable on scope, GPT on the gap. The WP1 methodology asks for market sizing, three margin stacks and a bounded range — not for an investment case. NPV and IRR belong to WP5, which scores entry mode. GPT is applying a criterion the mandate does not set for this work package. But GPT is right that a Board asked to allocate capital will want it, and right that the pilot ask has no number attached. Both models independently say: scope and cost the pilot.
Both name slides 5 and 16 as prose-heavy. They differ on the
separators: GPT calls slides 3, 7, 10 and 17 "sparse" and slide
17 "effectively an empty placeholder"; Fable treats the
separators as template furniture. Measured: slide 17
uses the Fin layout, whose master provides Nom / Rôle /
Email placeholders that the slide does not instantiate — so nothing
renders empty, but the closing page carries only "WP1 — Market"
with no doc ID, confidentiality line or leave-behind. Fable's fix is the
better one: "finish slide 17 with the doc ID, confidentiality line
and the three data-room requests as a leave-behind recap."
GPT scores a standards breach: "Slides 3, 7, 10 and 17 have no source line despite the no-exception rule." Fable and the L2 audit both treat dividers and book-ends as correctly exempt. The firm's standard says "Every page carries its source line. No exception." This is a consultant call on the firm's own rule, not a defect, and it should be settled once at cabinet level rather than re-argued each deck.
| Finding | Verification |
|---|---|
| Slides 15 and 16 claim the lease fee is "the only" Verdal figure not supplied — and both pages contradict themselves. Slide 15 item 2: "the only number in this analysis that Verdal has and we do not"; slide 16: "the only Verdal figure we had to invent". | Confirmed. The same slide 15 lists item 1 (convertible bays, "Held by Verdal today") and item 4 (energy contract, "Held by Verdal today") as also held and unsupplied. Three, not one. A factual self-contradiction on the two most action-oriented pages. |
tab 5!B31 contribution = €0.37/kWh (gross of
payment costs) while tab 5!B6 = €0.34 (net). Slide
13's "Contribution/kWh €0.25 to €0.45" does not say which definition it
varies. |
Confirmed. Two contribution definitions live on the same tab. |
| Slide 6's displayed steps do not tie to each other: 20,356 − 19,590 = 766; 766 − 718 = 48, against €47.6m in the model. Rounding applied per bar rather than from unrounded cells. | Confirmed. Same defect class as slide 8's percentages. |
| Slide 12's three largest cost lines (€48.2m of €70m) rest on a single adviser source — Interpath, May 2026 — with no independent second source. | Confirmed, and L1 adds that Interpath's figures are its own chart-footnote modelling assumptions, not operator disclosures. Interpath's own footnote assumes €180,000 per charger; the model uses €200,000. |
| The 5.8% reconciliation is partly circular — both routes share the same unsourced throughput assumptions, so "agreement" is weaker evidence than it looks. | Fair. tab 2!D17 says as much in the
model; the slide does not. |
| Finding | Verification |
|---|---|
| Possible double count between the Eurostat price and the grid capacity adder. Eurostat's non-household band IC price includes network charges; the model adds €0.04/kWh of "grid capacity and demand charges" on top. | Real risk, worth €3.37m a year on slide 12. Not a certain error — a large DC hub's medium-voltage capacity charge is billed on kW and is genuinely additional — but the model nowhere separates commodity, volumetric network, and fixed capacity charges. This needs one line of disclosure at minimum, and probably a decomposition. |
| Band mismatch worth checking. 84.315 GWh across 1,100 sites = 76.6 MWh per site per year, which is below the 500-2,000 MWh floor of Eurostat band IC. | Measured and confirmed. It is defensible if
charging shares the store's existing supply point (a supermarket's total
load sits comfortably in band IC or above) — but that is an assumption
the model does not state. One sentence in F24. |
| The €20.4bn plug-revenue figure applies an unweighted (DC + AC)/2 = €0.525/kWh to all European public charging, when public charging is DC-dominated by energy. | Confirmed. At a 70/30 DC/AC energy mix the figure is ~€21.5bn. Small in absolute terms; it undercuts the "precision" the number is presented with. |
| No page numbers on any slide. | Measured and confirmed — GPT is right and Fable is wrong here. The template carries no footer or slide-number placeholder on any of its four layouts. Slides 2-17 have no page number at all. (Fable read the "1" on slide 2 as a page number; it is the first support badge.) For a Board deck this is a real navigability defect — nobody can say "turn to page 12". |
| Verdal is not only a site host under self-operation — it also becomes the charge-point operator, so slide 4's "Verdal, today and under every structure tested" on the level-4 row is wrong. | Confirmed and conceptually clean. Under self-operation Verdal occupies levels 3 and 4. |
| Describing most plug revenue as energy-supplier pass-through is imprecise — the CPO retains the spread and pays only its delivered energy cost. | Fair. tab 1!F14 and slide 4 both use
the pass-through framing; it is a simplification a Board member in
energy will notice. |
Tier 1 — before the deck is shown to anyone (both models, plus L2)
tab 5, or strip those two
bars from slide 13. One or the other must happen.F37 and slide 11) and state the rate used.Tier 2 — before the interim on 8 September 7.
Put the confidence taxonomy on the slides. 8.
Give the lease-fee flip condition a visual on slide 11,
not a bullet. 9. Solve the break-even against the whole
estate (162.2, not 159.8) or re-label B9 as
DC-only. 10. Disclose or decompose the possible network-charge
double count, and state the Eurostat band basis. 11.
State that site eligibility has not been screened — one
line on slide 9, and a WP3 hand-off. 12. Scope and cost the
pilot so 8 September is a real budget decision. 13.
Re-date and version the deck for the interim. 14.
Add page numbers.
Tier 3 — improvements, not defects 15. Slide 5 as a table rather than three prose blocks, with currencies and horizon years visible. 16. Slide 2 from six supports to three or four. 17. End on a governing recommendation rather than a list of requests. 18. Correct slide 4's level-4 row: under self-operation Verdal is both operator and host.
Both models scored the same package and reached opposite delivery verdicts — 3.5/5 and "fixable in a day" against 2.2/5 and "HOLD". That divergence is not noise. It is the difference between judging the package against the WP1 methodology (Fable's frame, where the mandate asks for a sizing, three margin stacks and a bounded range, and gets them) and judging it against what a Board needs to allocate capital (GPT's frame, where the absence of a site screen, an investment case and a decision ask is disqualifying).
Both frames are legitimate and the choice between them is the consultant's. The honest reading is that Fable is right about WP1 and GPT is right about the engagement: the work package has done what it was scoped to do, and the deck currently presents it with more certainty than the model supports.
What is not in dispute: nine specific defects were found by both models independently, six of them also by the internal L2 pass. Those six should be fixed regardless of which frame the consultant adopts, and none of them requires new research.