Cookbook
View as Markdown

Call A/B test winners on a schedule

A weekly script that reads each campaign's A/B cohorts, applies the app's z-test and 50-lead floor, posts the verdict to Slack, and can promote the winner.

Last updated

An A/B test that nobody reads is a campaign sending half its leads the weaker copy. This recipe reads every running test once a week, applies the same statistics the app's Analytics tab uses (a pooled two-proportion z-test on positive-reply rate, with a floor of 50 contacted leads per arm before anything is called), posts the verdict to Slack, and, if you ask it to, promotes a significant winner so every new lead gets the better variation.

Before you start

  • An API key with campaigns:read, plus campaigns:write for --promote. See Authentication.
  • The date each test started, as --since. The cohort endpoint counts leads first contacted at or after it, so the arms are compared on the same population; one --since for all campaigns is fine when you started the tests together, otherwise run the script per campaign.
  • A Slack incoming webhook URL in SLACK_WEBHOOK_URL, or read the verdicts in the console.
  • Node.js 18 or later, or Python 3.10 or later.

How it works

  1. GET /v1/campaigns lists campaigns; GET /v1/campaigns/{campaign_id} says which have ab_testing_enabled.
  2. GET /v1/campaigns/{campaign_id}/ab-cohorts with since returns arms.a and arms.b, each with contacted, replied, positive and negative over leads first contacted since the test began.
  3. The script compares positive-reply rate with a pooled two-proportion z-test. Under 50 contacted leads in either arm the verdict is not_enough_data; a two-tailed p below 0.05 is a_wins or b_wins; below 0.2 is leaning_a or leaning_b; otherwise tie. These are the thresholds in the app's ab-stats and the API server's ab_stats.py, so the script and the Analytics tab agree.
  4. With --promote, a clear winner is promoted with POST /v1/campaigns/{campaign_id}/ab-testing/promote: variation A becomes the winner's copy, variation B is dropped, and the test ends. Leaning verdicts are reported, never promoted.

The script

// ab-verdict.mjs
const API_URL = "https://api.versionseven.ai/v1";
const API_KEY = process.env.VICTORIA_API_KEY;
const SLACK_WEBHOOK_URL = process.env.SLACK_WEBHOOK_URL;
const args = process.argv.slice(2);
const PROMOTE = args.includes("--promote");
const since = args[args.indexOf("--since") + 1];
if (!args.includes("--since") || !since) throw new Error("Usage: node ab-verdict.mjs --since 2026-09-01T00:00:00Z [--promote]");

// The same thresholds as the app's Analytics tab.
const MIN_ARM_SAMPLE = 50;
const SIGNIFICANT_P = 0.05;
const LEANING_P = 0.2;

const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

async function call(method, path, body) {
  for (let attempt = 1; ; attempt++) {
    const response = await fetch(`${API_URL}${path}`, {
      method,
      headers: { Authorization: `Bearer ${API_KEY}`, "Content-Type": "application/json" },
      body: body === undefined ? undefined : JSON.stringify(body),
    });
    const data = await response.json().catch(() => ({}));
    if ((response.status !== 429 && response.status !== 503) || attempt === 5) return { status: response.status, data };
    await sleep((Number(response.headers.get("Retry-After")) || 2 ** attempt) * 1000);
  }
}

// Abramowitz & Stegun 7.1.26, the same polynomial the app uses, so p-values round the same.
function erf(x) {
  const sign = x < 0 ? -1 : 1;
  const ax = Math.abs(x);
  const t = 1 / (1 + 0.3275911 * ax);
  const y = 1 - ((((1.061405429 * t - 1.453152027) * t + 1.421413741) * t - 0.284496736) * t + 0.254829592) * t * Math.exp(-ax * ax);
  return sign * y;
}
const normalCdf = (z) => 0.5 * (1 + erf(z / Math.SQRT2));

// Positive-reply rate of A vs B: {rateA, rateB, diff, z, p} with a two-tailed p.
function zTest(aN, aK, bN, bK) {
  const rateA = aN > 0 ? aK / aN : 0;
  const rateB = bN > 0 ? bK / bN : 0;
  const diff = rateA - rateB;
  if (aN === 0 || bN === 0) return { rateA, rateB, diff, z: 0, p: 1 };
  const pooled = (aK + bK) / (aN + bN);
  const se = Math.sqrt(pooled * (1 - pooled) * (1 / aN + 1 / bN));
  if (se === 0) return { rateA, rateB, diff, z: 0, p: 1 };
  const z = diff / se;
  return { rateA, rateB, diff, z, p: 2 * (1 - normalCdf(Math.abs(z))) };
}

function verdict(arms) {
  const a = arms.a ?? {};
  const b = arms.b ?? {};
  const test = zTest(a.contacted ?? 0, a.positive ?? 0, b.contacted ?? 0, b.positive ?? 0);
  let call = "tie";
  if ((a.contacted ?? 0) < MIN_ARM_SAMPLE || (b.contacted ?? 0) < MIN_ARM_SAMPLE) call = "not_enough_data";
  else if (test.diff !== 0 && test.p < SIGNIFICANT_P) call = test.diff > 0 ? "a_wins" : "b_wins";
  else if (test.diff !== 0 && test.p < LEANING_P) call = test.diff > 0 ? "leaning_a" : "leaning_b";
  return { call, test };
}

const pct = (rate) => `${(rate * 100).toFixed(1)}%`;
const lines = [];
const { data: list } = await call("GET", "/campaigns?limit=500");
for (const summary of list.campaigns) {
  const { data: detail } = await call("GET", `/campaigns/${summary.id}`);
  if (!detail.campaign?.ab_testing_enabled) continue;
  const { status, data } = await call("GET", `/campaigns/${summary.id}/ab-cohorts?since=${encodeURIComponent(since)}`);
  if (status !== 200) {
    lines.push(`*${summary.name}*: cohorts answered ${status} ${data.error}`);
    continue;
  }
  const { call: result, test } = verdict(data.arms);
  const a = data.arms.a ?? {};
  const b = data.arms.b ?? {};
  let line =
    `*${summary.name}*: ${result} · A ${a.positive ?? 0}/${a.contacted ?? 0} positive (${pct(test.rateA)}), ` +
    `B ${b.positive ?? 0}/${b.contacted ?? 0} (${pct(test.rateB)}), p = ${test.p.toFixed(3)}`;
  if (PROMOTE && (result === "a_wins" || result === "b_wins")) {
    const winner = result === "a_wins" ? "a" : "b";
    const promoted = await call("POST", `/campaigns/${summary.id}/ab-testing/promote`, { winner });
    line += promoted.status === 200 ? ` → promoted ${winner.toUpperCase()}; the test is over` : ` → promote failed: ${promoted.status} ${promoted.data.error}`;
  }
  lines.push(line);
  await sleep(650);
}

const text = lines.length ? [`:bar_chart: A/B verdicts since ${since}`, ...lines].join("\n") : "No campaign has an A/B test running.";
if (SLACK_WEBHOOK_URL) {
  const response = await fetch(SLACK_WEBHOOK_URL, { method: "POST", headers: { "Content-Type": "application/json" }, body: JSON.stringify({ text }) });
  if (!response.ok) throw new Error(`Slack answered ${response.status}`);
}
console.log(text);
node ab-verdict.mjs --since 2026-09-01T00:00:00Z
node ab-verdict.mjs --since 2026-09-01T00:00:00Z --promote
python ab_verdict.py --since 2026-09-01T00:00:00Z

A verdict line reads *Q3 outbound*: a_wins · A 18/120 positive (15.0%), B 6/115 (5.2%), p = 0.013. Weekly is the right cadence: the floor of 50 leads per arm takes most campaigns a week or two to reach, and reading a test daily invites calling it early.

What to expect

  • The floor is deliberate. With small arms a 1-of-10 against 0-of-9 split looks like a winner and isn't, so under 50 contacted leads in either arm the verdict is not_enough_data whatever the rates.
  • Cohorts, not totals. ab-cohorts counts only leads first contacted since since, with each lead in the arm it was assigned, so a lead contacted before the test began doesn't tilt it. GET /campaigns/{campaign_id}/analysis with variation=a or b reports the same arms over the whole campaign, and a connected assistant's campaign_analysis tool returns the verdict ready-made as ab_verdict.
  • Promotion is final. promote copies the winner into variation A, removes variation B and turns the test off; the campaign's traffic_split goes to 100. There's no undo beyond rewriting the sequence, which is why leaning verdicts are only reported.
  • A tie is information. Two variations that perform the same after 200 leads each means the thing you changed didn't matter to these buyers; test something bigger next.

Next steps