Category
AI
Last updated
August 2026
Related
Shopify integration

Stress test your AI stack: five questions, three setups

Five questions every DTC brand needs answered, run through point-solution MCPs, CSV exports in Claude, and the Polar MCP. Point solutions answered one, CSV two, Polar five.

Most brands build their AI stack backwards. They start with the tools: which MCP, which LLM, which context layer. Then, three weeks in, they find out the setup cannot answer the questions that actually decide budget. Start from the output instead. Write down the questions your business needs answered, then pick the stack that can deliver them.

Quick refresher first. Claude is the assistant; MCP, the Model Context Protocol, is how it reads your tools. Shopify ships an MCP, Klaviyo ships one, Meta Ads and Google Ads connect through their own MCP servers, and Polar ships the official Polar MCP. Connect them in Claude and you can ask your stack questions in plain English.

Here are the five questions we picked, one per function, because these are the ones that move money:

  • Merchandising. What were our top products by net profit last month?
  • Operations. Which orders are currently unfulfilled, and why?
  • Growth. What is our true ROAS by channel?
  • Retention. Which acquisition channel brings our lowest-LTV customers?
  • Finance. Does our P&L reconcile to Shopify gross?

Each one was run through three setups: point-solution MCPs (Shopify, Klaviyo, the ad platforms), CSV exports dropped into Claude, and the Polar MCP. Not because any of these tools is weak, but because a stack is only worth what it can answer. Below is the test, run end to end on a ~$31M/yr DTC brand. Then how to run it on yours.

The worked example, on a ~$31M/yr brand

Same model brand as the rest of this series: ~$31M/yr DTC, Shopify plus Klaviyo plus Meta and Google, reconciled in Polar. Every question below was asked three times, once through the single-source MCPs, once with CSV exports pasted into Claude, once through the Polar MCP. Here is the scorecard:

Scorecard
Five questions, three setups
QuestionPoint-solution MCPsCSV in ClaudePolar MCP
MerchandisingTop products by net profit last month
OperationsWhich orders are unfulfilled, and why
GrowthTrue ROAS by channel
RetentionWhich channel brings the lowest-LTV customers
FinanceDoes the P&L reconcile to Shopify gross
Point-solution MCPs 1/5  ·  CSV exports 2/5  ·  Polar MCP 5/5

The read: point-solution MCPs answered one of five. CSV exports answered two. The Polar MCP answered five. And the failures came from three different reasons, which is the interesting part. No single-tool MCP can answer a question that spans sources, and the questions that decide budget all do. Point several of them at Claude at once and the context overloads, so the output turns wrong rather than absent. CSV exports hand Claude raw rows with no definitions, so the moment you rebuild LTV from an order table and join it to channels, the odds of a wrong number go up roughly tenfold.

One honest note on this example: the brand is a model built on illustrative synthetic data, because we do not publish client numbers. The boundary it maps is real: run the same five questions on your own stack and you will hit it in the same places.

What you are building

A three-way stress test: five real operator questions, asked through your point-solution MCPs, through CSV exports in Claude, and through Claude reading Polar's governed model, with the boundary made visible. One question your store tools genuinely own. Four that reach across channels, margin, identity and finance, where a single-source setup hits the edge of its data and Polar keeps going. The output is an honest map of which job each tool is for.

What you need

  • Claude (Desktop or Cowork) with connectors enabled.
  • Your existing MCPs for the point-solution side of the test: Shopify, Klaviyo, and your ad platforms through their own MCP servers.
  • For the CSV side: an order export from Shopify, a COGS sheet, spend exports from each ad platform, and your P&L, dropped into a single Claude conversation.
Five raw CSV exports attached in a single Claude conversation
  • Shopify connected to Polar (native connector), ad platforms, Klaviyo, and your finance sheet connected, and the Polar Pixel installed.
  • The Polar MCP connected to Claude, for the Polar side of the test.

Question 1. Merchandising: what were our top products by net profit last month?

Ask the Shopify MCP and you get a clean ranking by revenue in seconds. That part lives entirely in the store and works well. But net profit per product needs COGS, shipping, payment fees, returns, and the ad spend attributed to each product, and none of that is in Shopify. The store MCP returns a confident answer to a different question, which is the failure mode worth knowing about.

CSV gets there, with effort: export orders, drop a COGS sheet next to it, and Claude can do the arithmetic. It is the one margin question where the join stays at the SKU level and the numbers hold.

Now connect the Polar MCP and ask the exact same question. The ranking comes back by net profit this time, and it is not the same ranking.

On the model brand, the difference is the whole point: the revenue number one, Hydra Serum at $412K, ranks third by net profit at $87K, while Collagen Peptides at $301K of revenue is the real winner at $126K. Rank the reorder on revenue and you celebrate the wrong SKU.

Question 2. Operations: which orders are currently unfulfilled, and why?

This is the one the store MCP wins outright, and honestly it is the better tool here. Fulfillment status, locations, stock levels, and the fix all live inside Shopify, so ask there and act there. CSV answers it too, but on a snapshot that was already stale when you exported it, and you cannot act from a spreadsheet. Credit where due: for in-store questions, the built-in path is right there and it is good. A stress test that pretends otherwise is not honest.

The Shopify MCP returning 214 unfulfilled orders split by reason

Question 3. Growth: what is our true ROAS by channel?

Here the boundary appears. True ROAS needs Meta, Google, and TikTok spend next to attributed Shopify revenue. That spend is not in Shopify, and each ad platform MCP only sees its own walled garden, so no single tool can do the division. You can ask Claude to stitch the MCPs together in the chat, but then it is inferring how numbers that were never designed to reconcile should join, and a subtly wrong join reads exactly like a right one. CSV has the same problem with worse inputs: each platform reports its own conversions on its own window, so summing the exports double-counts every order two platforms both claim.

Claude explaining why the ad platform MCPs cannot produce a trustworthy ROAS by channel

On Polar the spend is unified and revenue is pixel-attributed on one governed model: Meta 2.4, Google 2.9, blended paid 2.6 on the model brand. One question, one number, same number every run.

True ROAS by channel through the Polar MCP: Google 2.9, Meta 2.4, blended paid 2.6

Question 4. Retention: which acquisition channel brings our lowest-LTV customers?

This needs cross-channel attribution and full-history LTV, and neither lives in any single tool. Shopify knows orders, Klaviyo knows profiles, and each ad platform claims its own conversions; nobody holds the joined picture. The CSV route breaks down most clearly here: the order export carries an email but no acquisition source, the ad exports carry no customer identity at all, and one month of orders cannot show repeat behaviour. There is no key to join on. Asked plainly, Claude names the three things the answer would need and finds none of them in the files. Asked less carefully, that same missing join is exactly where a confident wrong number comes from.

Claude showing that LTV by acquisition channel cannot be rebuilt from raw CSV exports

Polar's pixel, identity resolution, and cohorts answer it directly: on the model brand, 90-day LTV runs $86 for Meta, $112 for Google, $131 for organic and referral. Meta is the lowest-LTV channel by a wide margin. It buys volume; organic buys value. That is a budget-shaping fact you cannot see from inside one tool, and it does not mean cut Meta, it means price it for what it actually returns.

90-day LTV by acquisition channel through the Polar MCP: Meta $86, Google $112, organic $131

Question 5. Finance: does our P&L reconcile to Shopify gross?

This is the question that ends the monthly argument between finance and growth, and no store tool can touch it. The Shopify MCP has no P&L. CSV gets you most of the way and then stops: the store knows returns, discounts and gift cards, so those three reconcile, but the last adjustment lives only in the finance sheet. You end up able to name the gap without being able to close it.

The P&L reconciliation stopping at a $38K unexplained residual without a unified model

On the model brand, Shopify gross sales for the month read $2.61M while the P&L revenue line reads $2.32M. The $294K gap is not an error, it is four definitional adjustments: $164K of returns, $121K of discounts, $47K of gift cards issued that are a liability rather than revenue, and $38K of shipping income the accountant books as revenue and Shopify does not. Polar holds the Shopify orders, the payment processor, and the finance sheet on the same model, so the reconciliation is a query that returns gross, each adjustment, and the P&L figure, with the bridge between them.

The full bridge from Shopify gross sales to the P&L revenue line through the Polar MCP

What the test shows

The boundary is not quality, it is scope and architecture. Scope: each MCP sees its own tool, so four of the five questions have no data to stand on, no matter how good the assistant reading it is. Architecture: point several single-tool MCPs at the same question and the context overloads, tools that generate a query on the fly infer how tables join and return a metric that is subtly wrong, and CSV hands Claude raw rows with no definitions at all. Polar sits on every source instead of one and builds a global ontology where each metric is defined once, deduplicated and attributed, so Claude reasons on the right data and the answer is the same every run.

Use the store MCP to run the store. Bring in Polar the moment the question reaches across the business.

Polar upgrade

One MCP that carries the whole business

Not optional for four of the five. The inputs live outside any single tool.

The exact thing Polar fixes: four of the five questions require data that lives outside any one tool, unified and defined once. Polar joins Shopify to 45+ sources on a governed model, so the cross-channel, attributed, margin and finance questions become answerable, and answerable consistently.

The honest limitation without it: there is no way to answer them from one tool's data, because the inputs are not there. That is a boundary of the data, not a shortcoming of any assistant working inside it. And neither workaround closes the gap: stitching MCPs in the chat makes the AI infer how sources join until the context overloads, and CSV exports hand it raw rows to guess with. A global ontology defines each metric once, deduplicated and attributed, so the answer is deterministic.

Connect the Polar MCP and rerun the five questions:

Prompt
Answer the five stress test questions using the Polar MCP: top products by net profit last month, unfulfilled orders and why, true ROAS by channel, the acquisition channel with the lowest 90-day LTV, and whether my P&L reconciles to Shopify gross. For each answer, name the sources it crosses.

With Polar the four unanswerable questions stop being unanswerable, and every number comes from a metric defined once in the semantic layer, so the granular cuts, per SKU, per channel, per customer, stay consistent on every run.

Polar Analytics

Starter prompts to extend it

  • Ask true ROAS, true CAC, and MER by channel, and note which of these a point-solution MCP cannot reach.
  • Which acquisition channel brings my lowest 90-day LTV customers, and which sources does the answer cross?
  • Bridge my Shopify gross sales to my P&L revenue line for last month, adjustment by adjustment.
  • Compute true contribution margin after ad spend and returns for my top collection.
  • Rerun the cross-channel questions across Polar's attribution models and show how the answer shifts.

A few honest notes

  • Right tool for the store job. For in-store questions, orders, fulfillment, inventory, the store MCP is the best tool, and acting on the answer happens there. This test is about scope, not quality.
  • The four cross-source questions are a boundary of the data, not of any assistant. No tool can compute with inputs it does not hold.
  • CSV is not a shortcut. It looks like it removes the integration problem, and it moves it into the model, where the wrong answer arrives with the same confidence as the right one.
  • Read-side. Everything in this test reads. Polar does not write to Shopify or Klaviyo; you act in the tools themselves.

Write your five questions before you pick anything. If a stack cannot answer them, it does not matter how good the demo was.