Ahlchemy Fulcrum scores your organization across four dimensions, pinpoints the one constraint holding you back, and names the best-fit AI engine and model — vendor-neutral, security- and governance-first.
A four-part diagnostic for the opening engagement. Work through it live as the client talks — it scores each leg, finds the binding constraint, and resolves to an honest starting point with a phased roadmap.
A discovery instrument you run live, during the call. You score what you hear across four legs; it finds the one weakest leg — the binding constraint — and resolves an honest starting point with a phased roadmap you can leave behind.
Companies want to start their AI journey at the exciting use case. The discipline this tool sells is starting at the constraint, not the ambition — the right first move is a function of the weakest of the four legs, which is usually driver clarity, the data foundation or governance, not the fun part.
Each leg meters 0–100 as you answer. The binding constraint is the lowest leg — that's what the recommendation keys on. It resolves to one of four starting points, in priority order:
Two flags can fire alongside any result: a guardrail flag when AI access is ungoverned (a sanctioned path belongs in Phase 0), and a capability flag when the team is the weak leg (run a guided pilot so the win survives your exit).
Resolve Starting Point generates the recommendation, reasoning, phased roadmap, and board-ready success criteria. Print / Save PDF produces a clean client leave-behind (it hides the controls and unselected options). New Assessment clears everything for the next call. Everything autosaves to this browser between sessions.
Generation became free. Judgment didn't. The projects that fail rarely fail on the model — they fail on the operating system around it: integration, data, governance, and someone deciding what's actually worth shipping. These are real patterns, anonymized. Each one is a leg of the diagnostic, skipped.
of enterprise generative-AI pilots delivered no measurable P&L impact.
The 5% that created real value didn't have better models. They had better operations — integration, clean data, and judgment. AI isn't a model you buy; it's a system you operate.
Roughly 40% of companies have an official, sanctioned AI subscription — but around 90% of employees use their own AI for work anyway. Half your organization is already building in the shadows, often with customer data, in tools you don't control.
of U.S. internet households now use AI tools — up from 51% a year earlier. The capability is everywhere; the governance almost never keeps pace with it. Parks Associates, 2026
A ~$14B fintech pointed a raw model straight at its customers, cut average resolution time from 11 minutes to 2 across 2.3M chats, and projected ~$40M in savings by replacing its support agents. The quarter's dashboards looked spectacular.
Why it stalledWithin a year it was quietly rehiring humans. The average was fine; the tails weren't — the edge cases, the angry customer, the moment that needed accountability. The metric that mattered (trust) never showed up on the dashboard until it was gone.
An executive asked an AI for a command-center dashboard. It arrived in 41 seconds — trend lines, a "readiness score" of 94.2, a "digital maturity index." It looked like a quarter of work from a BI team.
Why it stalledIt solved no stated problem. Nobody had asked for it. No decision changed because of it. The headline metrics were invented by the model. Polish used to be expensive — that's why we trusted it. Now polish is free, and it's the easiest thing in the world to mistake for substance.
The same model got the same word-for-word prompt — "write our Q3 executive strategy" — twice. Once against clean, current data; once against the company's actual operational data.
Why it stalledClean data produced a sound, defensible plan. The real data — duplicate customer records merged into one, a refund policy pulled from a stale 2019 file, a half-finished field — produced a confident, dangerous one: consolidate three "different" customers who were the same person, and "eliminate any customer who touches the red flag." Same model. The variable was never the model.
A company with no approved AI path assumed that meant no AI risk. In reality, most of the staff were already pasting work — contracts, customer lists, support transcripts — into consumer chatbots on personal accounts to hit their numbers.
Why it stalledYou can't govern what you've refused to acknowledge. Banning AI doesn't stop adoption; it just pushes it into tools with no logging, no DLP, and no contract protecting your data. The first incident is a data-exposure one, and it ends the whole program.
They were operating problems — a binding constraint named too late, or never. That's the entire point of the diagnostic: find the one leg that's holding everything back before you spend a quarter proving it the hard way.
This is a complete assessment for a fictional company, Meridian Retail Group, so you can see exactly what the engagement produces. The top is shown in full; the starting-point detail, phased roadmap, governance mapping and engine recommendation are what we deliver in the paid engagement.
The driver is genuine and the data foundation is coming along (foundation 61/100), but AI is already in the building through ungoverned consumer tools — customer and order data is walking out the door into accounts you don't control (governance 32/100). This is the most urgent problem even though it isn't the one leadership asked about. You don't have an adoption problem yet; you have a governance problem, and the fix is a safe, approved way to use AI before a single flagship use case ships.
Controls mapped to the frameworks in scope — NIST AI RMF (Govern / Map / Measure / Manage), the EU AI Act risk tiers, and ISO 42001 — with PCI-DSS and CCPA/CPRA obligations called out where customer data meets model access.
The personalization engine the business wants is high value but low feasibility today — sequence the governance enablers first. A governed internal-support assistant is the easy, high-confidence first win once the sanctioned path exists.
For a PCI/CCPA-bound retailer, route access through a cloud platform under your existing contract and region rather than a direct consumer app, so data handling, logging and DLP are enforced by default. Match the model tier to the task — a fast, lower-cost tier for high-volume support classification, a frontier tier only where reasoning genuinely demands it.
The complete starting point, the phased roadmap, governance mapping and the engine & model recommendation are delivered as part of your engagement.
The use case comes first; the engine and model come second. This is a vendor-neutral map of the major GenAI platforms and access routes — a decision framework to pick the best fit for a customer, not a leaderboard.
Picking the right AI, ML, or generative engine matters — but it's equally critical to put the right security and governance around it, and to build real skill in prompting. Prompting is a craft: it's developed through use as much as through training. Budget for hands-on practice, internal patterns, and review — not just a tool-selection decision. A capable team on a governed platform will out-perform a better model used carelessly.
Prompting is the highest-leverage skill in an AI program and the one most teams underinvest in. These patterns hold across every engine and model — they're how you get consistent, reviewable output instead of one-off luck.
The strongest results often come from a chain, not a single call: each model does what it's best at and hands a clean, structured artifact to the next. This is a core capability to build — and to govern — not an afterthought.
Make the handoff reliable: pass a contract (the exact fields the next step expects), not loose prose, and validate it before handing off; keep a human or an automated check at the seam, because chained errors compound; and govern the whole chain — data handling, logging and cost apply to every hop, not just the first. Match the model to the step: don't pay frontier prices for extraction, and don't hand final synthesis to a tiny model.
What's new in the Ahlchemy Fulcrum assessment.
?mode=demo).Everyone who requested the full assessment, plus every engagement you've saved to the cloud. These land in your D1 database — nothing is emailed anywhere, so this is where you follow up.
Run the full tool yourself from here — or reopen a lead from the list above with "Open".
Your offline study material lives on its own page — use the Study tab, or:
Install Fulcrum as an app for true offline use; it opens in its own window and the pages work without a connection.
iPhone/iPad (Safari): Share → Add to Home Screen. Desktop Chrome/Edge: the install icon in the address bar, or the button above.
Your offline reference for client calls and the stage. Everything here loads with the page, so it works without a connection once the app has opened (install it from the Admin page for a true offline app).
What the terms mean, how the major models differ, how you actually reach them, and how to pair a model and a toolset to a job. Written to make you fluent on a client call, not to pass an exam.
Rule of thumb for clients: if the job is "predict/score from structured data," it's probably classic ML. If it's "understand or generate language/content," it's GenAI.
| Family | Character & strengths | How you reach it | Reach for it when… |
|---|---|---|---|
| Anthropic — Claude | Strong reasoning, long-document work, careful instruction-following; safety-forward. Clear "no training on your API data" posture. | Direct API; AWS Bedrock; Google Vertex; Claude Team/Enterprise | Regulated or document-heavy work; teams on AWS/GCP. |
| OpenAI — GPT / ChatGPT | Broadest ecosystem and tooling; strong general + multimodal; fast feature cadence. | Direct API; Azure OpenAI; ChatGPT Enterprise | Widest integrations, off-the-shelf productivity, Azure shops. |
| Google — Gemini | Deep Google Cloud / Workspace integration; strong multimodal and long context. | Direct API; Google Vertex AI; Gemini for Workspace | Google-stack orgs; heavy multimodal or BigQuery data. |
| Meta — Llama (open weight) | Download-and-run models; full control and fine-tuning; no per-token fee. | Self-host; or hosted on Bedrock/Vertex/others | Data sovereignty, air-gapped, or cost-at-scale with ML talent. |
| Mistral (open + API) | Efficient open-weight and API models; EU-domiciled. | Direct API; self-host; cloud platforms | EU residency, cost/latency-sensitive, self-host. |
| xAI — Grok | Fast-moving; real-time signal via X; strong on some reasoning. | Direct API; X ecosystem | Real-time/social signal; speed over enterprise maturity. |
| Specialized models | Embeddings (for search/RAG), image (diffusion), speech-to-text & TTS, video, re-rankers — often from the same vendors or open models. | Same APIs / open weights | You need a specific modality, not a general chat model. |
Vendor-neutral truth: on mainstream tasks the frontier models are close. The differentiator for a client is usually the access route and data posture, not a benchmark point.
"The model" and "how you call it" are different decisions. Choose the route for compliance first, then the model for the task.
API literacy to carry: API key (secret, server-side only), tokens (what you're billed on), rate limits (requests/tokens per minute), streaming (tokens arrive live), context window (input cap), and data posture (is your data used for training? on enterprise routes, generally no — confirm per vendor).
| Use case | What it really needs | Model tier | Supporting tools |
|---|---|---|---|
| Customer / internal support assistant | Grounding in your content; a human seam on risky replies | Mid; small for triage | RAG + vector store, guardrails/PII redaction, logging |
| Coding assistant | Strong code reasoning; repo context | Frontier/mid | IDE copilot, code search, test harness, human review |
| Document Q&A / knowledge | Retrieve the right passages, cite them | Mid | Embeddings + vector DB (RAG), re-ranker, citations |
| High-volume extraction / classification | Cheap, consistent, structured output | Small / fast | JSON/schema output, eval set, batch processing |
| Summarization | Enough context window; faithful to source | Small–mid | Chunking, grounding, a factuality check |
| Agents / multi-step workflows | Planning + reliable tool calls | Frontier/reasoning | Function calling, orchestration framework, tracing, limits |
| Image generation | A diffusion model, not an LLM | Specialized (image) | Prompt patterns, brand/style guardrails, review |
| Transcription / voice | Speech-to-text (and TTS to speak) | Specialized (speech) | Diarization, redaction, an LLM to summarize after |
| Semantic search | Meaning-based matching, not keywords | Embeddings model | Vector DB, hybrid (keyword+vector), re-ranking |
The pattern: pick the capability the job needs, choose the smallest tier that clears the bar, and wrap it with the retrieval, structure and review it needs to be trustworthy.
Prompting is the highest-leverage skill and it compounds with use. Set role, goal and audience; show examples; supply the context; ask for the exact format; constrain and verify; keep a shared prompt library. And the strongest results often come from a chain — one model (or model + tool) hands a clean, structured artifact to the next, e.g. a reasoning model writing a spec that an image model renders.
Full detail lives on the Solutions page ("Prompting that compounds" and "Orchestration — using one AI to feed another").
The whole talk, beat by beat, with the numbers to memorize, the lines to land, and the questions you'll get. Built to study cold on a plane.
The promise to the room: you'll leave able to tell the difference between AI that changes a decision and AI that's just theater — and you'll know the one move that every version of the future rewards: judgment.
Who's in the room: executives and operators who are equal parts excited and afraid. Half have an approved AI tool; nearly all are already using one anyway. Meet the fear honestly, then reframe it.
Verify any figure before you quote it live — studies get revised. If challenged on a number, cite the source and move on; don't defend a decimal.
The studies, frameworks, and real-world cases behind the diagnostic — with a starting path. Use these to back a claim, pick a framework, or arm a story on a client call.
Tap any item to drill into the detail.
The 95% stat — value comes from operations, not from a better model.
A 2025 study of real enterprise GenAI deployments (not lab benchmarks) measuring whether pilots reached production and moved the P&L.
Open with it to move the conversation from "which model do we buy" to "which operations do we fix." It's the evidence behind the whole diagnostic.
Stop evaluating models; invest in integration, data readiness and workflow redesign around one high-value process.
MIT NANDA / Project NANDA, "The GenAI Divide: State of AI in Business" (2025). Verify the exact sample and figure before quoting.
The neutral, most-cited baseline on capability, cost, adoption and investment.
A large annual, vendor-neutral report tracking the state of AI across research, technical performance, economy, policy, education and responsible AI.
When an exec distrusts vendor claims, cite the Index for a defensible, independent "state of the field" number.
Calibrate budget and expectations to real trend lines instead of hype or fear.
Stanford HAI, "AI Index Report" (latest annual edition).
Where enterprises actually capture value, by function — and where they don't.
An annual global survey of organizations on AI adoption, value capture, and practices.
Benchmark a client against peers by function and maturity; justify focusing the first use case where value actually lands.
Pick the first use case in a function with proven value capture; commit to workflow redesign, not a tool drop-in.
McKinsey Global Survey on AI (annual).
The default, vendor-neutral governance spine for a U.S. enterprise.
A voluntary framework for identifying and managing AI risks across the lifecycle, widely adopted as the U.S. governance backbone.
A 2024 companion that names GenAI-specific risks (hallucination, data leakage, IP, prompt injection, CBRN/dangerous content) and suggested actions.
Adopt the four functions as the client's governance operating model; structure every AI risk review around Govern/Map/Measure/Manage.
Make this the backbone of the governance leg; map controls to it rather than inventing an ad-hoc process.
NIST AI RMF 1.0 + the Generative AI Profile (NIST-AI-600-1).
Risk-tiered obligations for anyone touching EU customers or data.
The EU's horizontal, risk-based AI regulation — the first comprehensive AI law, with extraterritorial reach (it applies if your AI's output is used in the EU).
Phased: prohibitions first, GPAI duties, then high-risk obligations — rolling in across 2025–2027.
For any EU exposure, classify each use case's tier early; the tier drives the documentation and controls you owe.
Run a risk-tier classification before building; high-risk changes the whole plan (and cost).
Regulation (EU) 2024/1689. Confirm current enforcement dates.
The "ISO 9001 for AI" — a certifiable AI management system.
The first international management-system standard for AI (an "AIMS"), designed to be audited and certified like ISO 9001 or 27001.
Reach for it when a client needs a certifiable, auditable posture — regulated industries, enterprise procurement, or customers demanding assurance.
Decide whether to stand up a formal AIMS (and pursue certification) vs. a lighter NIST-based program.
ISO/IEC 42001:2023.
A large share of GenAI projects are abandoned after proof-of-concept.
Analyst projections on GenAI adoption, spend and failure/abandonment rates.
Gartner has projected that a significant share of GenAI projects (commonly cited around a third) are abandoned after PoC — driven by poor data quality, unclear business value, inadequate risk controls and escalating cost.
Create urgency to "fix one thing completely" and avoid pilot sprawl; it pairs with the MIT 95% figure.
Fund fewer, deeper initiatives with a clear value case and data foundation — not a portfolio of PoCs.
Gartner press releases / analyst notes — verify the current figure and date before quoting.
Tap any case for the full story, cause and lesson. Names are here for your study — anonymize on stage.
Air Canada's support chatbot told a grieving customer he could claim a bereavement discount retroactively — which wasn't the airline's policy. When the airline refused, a tribunal (2024) held it liable for what its bot said, rejecting the argument that the chatbot was "a separate legal entity."
A customer-facing bot was left to generate policy answers with no grounding in the real policy and no human seam on commitments.
Klarna announced in 2024 that its AI assistant was doing the work of ~700 agents and projected ~$40M in savings. By 2025 it walked it back — restoring human options and rehiring — citing quality and customer experience.
Deflection and average handle-time looked great; the edge cases — the upset customer, the unusual request — are where trust is won or lost, and those didn't show on the dashboard.
Zillow's "Offers" iBuying business used an automated pricing model to buy homes. In 2021 it shut the program down, took a write-down of roughly $300M+, and cut ~25% of staff after the model mispriced a turning market.
A forecast model was trusted to make large, hard-to-reverse financial bets; it inherited the limits of its data and assumptions when conditions changed.
Amazon built an experimental résumé-screening model and scrapped it (reported ~2018) after finding it penalized résumés containing "women's" and downgraded graduates of women's colleges.
It learned from ~10 years of historical hiring data that skewed male — so it faithfully reproduced the bias.
Samsung engineers (2023) pasted confidential source code and internal notes into a consumer AI chatbot to get help; the company then restricted employee use of such tools.
With no sanctioned, governed path, capable employees used the shadow one — sending proprietary data to a third party with no contract protecting it.
New York City's "MyCity" business chatbot (2024) confidently told business owners they could do illegal things — e.g., take workers' tips or fire someone for reporting harassment.
A high-stakes advice bot was shipped without grounding in authoritative rules, without review, and without a clear owner accountable for wrong answers.
Morgan Stanley, with OpenAI, built an assistant that answers advisors' questions from the firm's own vetted research library (a RAG system over owned, curated content). High adoption among advisors.
Internal-first (not pointed at clients), grounded in trusted content with citations, and a professional stays in the loop — it augments the human instead of replacing the relationship.
Controlled studies on AI coding assistants (e.g., a GitHub Copilot RCT) showed developers completing a scoped task meaningfully faster (~55% in one study) with a human still reviewing output.
Well-scoped, reversible work where the reviewer is the user who catches errors — speed where mistakes are cheap to undo.
Teams that used AI to draft agent replies (with a human reviewing and sending) captured real efficiency without the trust blow-up that full automation caused elsewhere.
The handoff is the design: the model does the heavy lifting; the human owns the last step and the accountability.
Enterprises that routed access through their cloud platform — region, logging, DLP, BAA — got broad workforce adoption and compliance.
The sanctioned path was made easier than the shadow one, so people used it; governance was built in rather than bolted on.
Named here for your study. On stage, name the pattern and the lesson; anonymize the company unless you're certain of the current facts.
Every recommendation the engine can land, rendered in full on a representative company, so you know the entire deliverable cold — the "why here," the discipline, the roadmap, and the board-ready success criteria for each.
The driver is a board mandate without a measurable outcome (driver clarity 34/100). Anything built now technically works and satisfies no one, because no one has agreed what "works" means. The first deliverable is agreement, not architecture.
Deliberately provisional — you lock this only after the driver is agreed. But for a private manufacturer under ITAR/EAR the shortlist is predictable, so pre-stage it:
| Layer | Recommendation | Why here |
|---|---|---|
| Engine route | In-tenant / sovereign — Azure OpenAI in a controlled enclave, or AWS Bedrock (GovCloud where export scope demands it) | ITAR/EAR means technical data can't transit uncontrolled infrastructure. Rules out consumer tools and shared SaaS. |
| Primary model | Claude (Sonnet-class via Bedrock) or GPT-4o-class via Azure OpenAI | Both run inside your cloud boundary with no training on your data. Pick by which cloud already holds the controlled enclave. |
| Runner-up | Llama-class, self-hosted | Fallback for anything that must be fully air-gapped. Lower ceiling, but nothing leaves the wire. |
| Dev tools | GitHub Copilot (enterprise, telemetry off); a thin orchestration layer (LangChain / Semantic Kernel) only if an app follows | Keep the toolchain inside already-cleared vendors. |
| APIs / integrations | Bedrock / Azure OpenAI SDK behind an internal gateway; identity via existing SSO | One governed seam so the model stays swappable and every call is logged. |
| Add-ons | Private retrieval (vector store inside the enclave); DLP on the gateway | Ground answers on controlled documents without those documents leaving the boundary. |
AI is already in the building through ungoverned consumer tools — customer and order data walking out the door (governance 30/100). This is the most urgent problem even though it isn't the one leadership asked about. You don't have an adoption problem yet; you have a governance problem, and the fix is a safe, approved way to use AI.
Controls mapped to NIST AI RMF (Govern / Map / Measure / Manage), the EU AI Act risk tiers, and ISO 42001 — with PCI-DSS and CCPA/CPRA obligations called out where customer data meets model access.
The deliverable here is the sanctioned path itself — the stack is the guardrail, not a feature. Make the approved route easier than the shadow one:
| Layer | Recommendation | Why here |
|---|---|---|
| Engine route | Identity-brokered gateway in front of enterprise model APIs (no-train tier) | Everyone reaches models through one logged, controlled door — not consumer apps that pocket customer data. |
| Primary model | Claude (enterprise) or GPT-4o (enterprise, no-train) | Contractual no-training + data-residency terms are the whole point under PCI/CCPA. |
| Runner-up | Azure OpenAI in-tenant | When a use case touches cardholder or regulated PII and must stay inside your cloud. |
| Dev tools | Copilot (enterprise, no-retain); approved prompt library published internally | Give people a blessed, convenient option so they stop pasting into random tools. |
| APIs / integrations | SSO/SCIM for access; gateway SDK; SIEM feed for every call | Access tied to identity means instant revoke and a real audit trail. |
| Add-ons | PII/PAN redaction at the gateway; DLP; policy-as-code guardrails | Strip cardholder data before it reaches a model; prove it to an auditor. |
The driver is real, but the foundation can't support it yet (foundation readiness 38/100). The honest conversation: you want the use case, but your data/platform layer isn't ready, so the first project is the boring enabler that makes the exciting one possible. This is exactly why they hire an outside voice who can say it.
Under HIPAA the foundation is the project — the model is the easy part. Everything below assumes a signed BAA and de-identification before model access:
| Layer | Recommendation | Why here |
|---|---|---|
| Engine route | HIPAA-eligible, BAA-covered — Azure OpenAI or AWS Bedrock under a Business Associate Agreement | No PHI may touch a model without a BAA. This is non-negotiable and it rules out consumer tools. |
| Primary model | Claude via Bedrock (BAA) or GPT-4o via Azure OpenAI (BAA) | Both operate inside the covered boundary with no training on your data. |
| Data platform | Governed lakehouse + a de-identification pipeline (Safe Harbor / Expert Determination) | The binding gap: data isn't accessible, clean, or de-identified yet. Fix this before any use case. |
| Dev tools | Access-controlled notebooks; data-contract tests; lineage tracking | Prove where every field came from and who can see it. |
| APIs / integrations | FHIR/HL7 connectors; identity via SSO; gateway with full audit | Meet the data where it lives in health systems, governed end to end. |
| Add-ons | De-id service; consent/authorization tracking; minimum-necessary filters | Enforce HIPAA's minimum-necessary rule in the pipeline, not by policy alone. |
Clear driver, ready foundation, capable team — the healthy case. Start with a deliberately small pilot: chosen so a win is legible to the board and a loss is survivable. The point isn't the pilot's output; it's the internal capability and confidence it builds.
Capability is the weakest leg here (65/100). Run the pilot as a guided one with capability-building baked in, or the win won't be sustained after the engagement ends.
The healthy case — ready to build. Go API-first and keep the model behind a swappable seam so you can chase the price/performance curve:
| Layer | Recommendation | Why here |
|---|---|---|
| Engine route | API-first — Anthropic / OpenAI direct (enterprise, no-train) or via Bedrock | SOC 2 is satisfiable with enterprise terms; direct APIs give you the newest models fastest. |
| Primary model | Claude Sonnet-class for the pilot's core reasoning | Strong quality-to-cost for production workloads; easy to swap up to a frontier tier for hard calls. |
| Runner-up | GPT-4o-class or Gemini-class behind the same seam | Keep two routes wired so a price or quality shift is a config change, not a rebuild. |
| Dev tools | Copilot / Cursor for the team; an eval harness from day one | Capability is the weak leg — tooling that teaches while it ships is the point. |
| APIs / integrations | Provider SDKs; LangGraph / lightweight orchestration; feature-flagged model router | One seam, instrumented, so ownership transfers cleanly to the internal champion. |
| Add-ons | Observability + eval (Langfuse / Braintrust); guardrails; prompt-version control | Measure quality against the target and keep it from regressing after you leave. |