00 The short version
How I measure results
I measure from inside the ad platforms and the CRM, not from a dashboard that averages everything into one flattering number. Before I judge a channel, I make sure the tracking underneath it is telling the truth, because a bid is only as good as the data it optimizes toward. Then I separate what the ads created from what they only harvested. And I do not claim a result I have not verified in the account itself.
01 First principle
Tracking quality
before optimization.
Bad data leads to bad bids. Nothing downstream matters if the numbers feeding it are wrong, so the measurement layer gets rebuilt first.
The data comes before the bid
Smart Bidding optimizes toward exactly what you feed it. If the conversions are miscounted, missing, or importing at zero dollars, every optimization after that compounds the error. I rebuild the tracking and confirm it before I touch a budget. The reference architectures I work from are the Tracking Stack for ecommerce and the Lead Quality Stack for lead-gen accounts.
A clean audit gates the spend
The engagement opens with a measurement audit, and I do not move a budget dollar until it lands clean. That order is not caution for its own sake. It is the only way the numbers I report afterward mean anything. The conversion tracking diagnostics walk through the signals that say a stack is lying.
Fixed in dependency order
The fixes are sequenced because the later ones depend on the earlier ones being right. Identity and event collection before destinations, destinations before values, values before bidding. Optimize on top of a broken layer and you optimize the wrong thing faster.
02 The honest limits
What attribution can
and cannot prove.
Attribution is a useful model, not a measurement of truth. Reading it as truth is how founders cut the wrong budget.
The platforms will never agree, and that is fine
GA4 sessionizes, each ad platform claims what it can see, and Shopify counts orders. Three systems, three methods, three numbers. I reconcile them within a roughly fifteen percent band rather than chase a false match, and I pick the right source per decision. The reconciliation logic is here.
Attributed is not incremental
Platform ROAS counts every sale the platform can claim, including demand the ads only harvested. It is not a measurement of what the ads added. Where the budget justifies it, an incrementality holdout answers the harder question, and the honest answer is almost always that the ads added less than the platform claimed. I wrote about this in what a 21x ROAS actually hides.
Blended numbers hide where the money comes from
A single account average blends demand creation and demand harvest into one figure that describes neither. I pull ROAS by campaign type, so the prospecting that grows the business is judged on its own math and not carried by a branded or Shopping campaign doing a different job.
03 The read
How I evaluate
performance.
The same read every week, in the same order, so the trend is legible and the decisions are documented.
Judged inside the platform, on the primary actions
I evaluate campaigns on the conversion actions bidding reads, confirmed as primary in the account, rather than on a downstream dashboard that reassembles the data its own way.
Lead-gen is judged on closed-won, not lead volume
A cheap lead that never signs is not a win. For service businesses I push the outcome back to the CRM and evaluate on closed cases and revenue, which is what the Lead Quality Stack exists to make possible.
A written summary every Monday
Spend, revenue, ROAS, new versus returning, the campaigns that moved, and the decisions made. Nothing important lives only in a call that already happened. The cadence is part of the Operator Method.
04 The bar for a claim
What I need before
I claim a result.
A result on this site is something I verified, not something a dashboard suggested. The case studies show the standard.
Verified read-only in the account
The figure is read from inside the account or the platform it lives in. When a scoring system makes a judgment, I benchmark it against labels a human already made before trusting it. That is how the law firm call-scoring build was checked, lead by lead, against the firm's own labels.
Measured now, or measured later, never asserted early
Some effects take weeks to compound. When that is true I claim the part that is verified today, like a restored signal or recovered event volume, and I measure the revenue effect in a later update instead of guessing it now. The server-side funnel case is written exactly this way. On that store I measured first-party routing recovering 31.7 percent of events and half of purchases that browser tracking prevention would otherwise have dropped.
Benchmarks that exclude the broken accounts
When I publish benchmark figures, accounts with broken tracking are excluded, so the numbers are not poisoned by data that was never real. You can see the method on the home services paid search benchmarks.
05 FAQ
The questions
I get on this.
- Why do GA4 and my ad platform never show the same conversions?
- Because they count different things. GA4 sessionizes and attributes on its own model, each ad platform claims the conversions it can see, and Shopify counts orders. They were never designed to agree. I do not try to force them to match. I reconcile them within a roughly fifteen percent band and pick the source of truth per decision: the store for revenue, the platform for in-auction bidding signal.
- Do you report platform ROAS or real ROAS?
- Platform ROAS is attribution, not proof of incremental return. It counts every sale the platform can claim, including demand the ads only harvested. I report it because bidding runs on it, but I read it against total account revenue and, where the budget justifies it, an incrementality holdout. A holdout tells you what the ads added, and it is usually less than the platform claims.
- Why fix tracking before optimizing the ads?
- Because a bid is only as good as the data it optimizes toward. If the conversions feeding Smart Bidding are miscounted, missing, or valued at zero, every optimization after that is built on a lie. I rebuild the measurement layer and confirm it in the account before I move a budget dollar. It is the least glamorous work and the highest leverage.
- What counts as proof of a result before you claim it?
- A figure verified read-only inside the account or the platform it lives in, not a number off a dashboard I have not checked. Where a human benchmark exists, I test my scoring against it first. And I separate what is measured now from what takes weeks to compound. If a result needs more time to be real, I say so and measure it later instead of asserting it early.
If your numbers do not reconcile
Fix the measurement.
Then trust the number.
If you are not sure your tracking is telling the truth, that is the first thing I check. A tracking audit answers it, and the free Setup Audit is where it starts.