Case study · Northwind Outfitters

A representative mid-market pilot

From missed moves to a governed ops floor in 90 days.

Northwind Outfitters — a fictional upper-mid-market apparel and home goods retailer with a 12,400-SKU catalog across three marketplaces and a 3-person operations team — rolled out the Merchanaut agent fleet over a single quarter. This page walks the baseline ops pain they were working with, the order the agents came up in, and the quantified outcomes over the first 90 days of live moves.

All numbers on this page are from a representative pilot scenario. Northwind Outfitters is a fictional retailer — the brief here is to show the kind of measurable shift the platform produces, not to cite a real customer.

The retailer

An upper-mid-market multi-channel floor.

Northwind Outfitters sits in the band Merchanaut is built for — past the point where spreadsheets and a single-channel plugin still work, short of the size where a full in-house ops team pays for itself. The pilot ran against a representative slice of the catalog and the same supplier list the ops team already buys from.

Channels

D2C site + 3 marketplaces

Operations team

3 people — ops lead, analyst, returns

GMV band (pilot)

~$45M annualised

Catalog size

~12,400 live SKUs across 8 categories

Baseline ops pain

Three symptoms of a team running past capacity.

Before any agent came up, the operations team was carrying the full ops load on a calendar that had stopped fitting the work. The three symptoms below are the ones the rollout had to address.

A 3-person ops team running at capacity.

Catalog updates, pricing reviews, restock calls, and refund rows were piled on the same calendar. The ops lead was the de-facto pricing analyst, the buyer, and the returns approver — and there was no separation of duties for any of them.

Mid-week competitor price moves shipping too late.

Competitor repricing on Monday through Thursday was caught on Friday in a manual scrub. By the time the catalog was updated, two of the most visible categories had either lost rank or softened margin to defend it.

A weekend inbox backlog that bled into Monday.

Customer messages queued from Friday close-of-business to Monday open averaged a 60-hour first-touch window. The returns queue sat even longer — refunds above the policy matrix carried over week to week.

Merchanaut rollout

A staggered, four-week-by-four-week enablement.

The agents came up in a deliberate order — read-only and dry-run first, live writes once the ops team had signed off on each lane, then the supplier and inbox loops once the catalog work was steady. Every move, at every stage, passed the same guardrail.

  1. Weeks 1–2

    Baseline, dry-run, and the first repricing passes.

    Inventory Sentinel wired into the WMS and shipped its first stock-out forecasts. Price Officer ran in dry-run against 8,000 SKUs, replaying the prior 60 days of competitor moves against the proposed guardrails and the ops team's existing pricing bands. Bob-Test blocked 211 of those moves at the guardrail — every block logged with a board-readable reason.

  2. Weeks 3–4

    Catalog Spy goes live. Price Officer cuts its first live band.

    Catalog Spy started capturing competitor deltas inside the retailers' pricing windows, with CFAA-compliant rate limits and CAPTCHA handling on by default. Price Officer cut its first live moves on the eight highest-velocity categories, inside the bands the ops team had signed off on, with Bob-Test repeating an FTC and Sherman check on every line.

  3. Weeks 5–8

    Order & Refunds Desk takes the order and returns queue.

    Order & Refunds Desk cleared the order queue end-to-end against the retailers' policies and processed returns inside the refund matrix. The small percentage that needed a human — disputed chargebacks, edge-case SKUs, sensitive-tone tickets — shipped to the ops lead with the policy row, the evidence, and a suggested reply attached.

  4. Weeks 9–12

    Restock Buyer closes the supplier loop. Campaign & Inbox opens the upsell loop.

    Restock Buyer drafted and sent first replenishment POs against the existing supplier catalogue with the negotiated terms the ops team already had on file. Campaign & Inbox fired segmentation-driven upsell campaigns in the retailers' tone of voice, with sentiment-aware escalation when a customer signal tilted the queue.

90-day outcomes

Four numbers the operations team reads on a single page.

The figures below are representative — taken from the kind of measurable shift the platform produces on a pilot of this shape, not from any specific customer. The full audit trail is replayable: same snapshot plus the same policy version produces the same decision.

+1.8 pts

Margin recovered

Across the eight categories the pilot ran against, gross margin moved from a flat 31.4% to a sustained 33.2% over a 90-day window. The lift was most pronounced on the mid-tail catalog where repricing velocity was previously bottlenecked by manual review.

36 hrs → <30 min

Repricing response time

A representative competitor price move went from a 36-hour first-pass lag to under 30 minutes, end-to-end including sign-off at the guardrail. The ops team moved from catch-up to forward-leaning.

42 of 58

Supplier counter-offers closed

A representative supplier-counters run landed 42 of 58 negotiated counter-offers inside target terms, with the remaining 16 escalating to a human buyer with the evidence already attached.

2 FTEs

Redeployed to category expansion

The analyst and the returns approver moved off the operational backlog and onto category expansion — the kinds of work that historically had no calendar slot behind the running queue.

After the rollout

The ops team now works the floor, not the queue.

With the agents running the recurring operational load, the analyst and the returns approver moved into category expansion and strategic sourcing — the work that has no calendar slot behind a running queue. The ops lead now reviews agent decisions on the audit page, edits guardrails on the policy surface, and walks the supplier counter-offers that need a human voice.

“The reason these results compound is that every agent move is governed, logged, and bounded. The team stopped firefighting the queue and started tuning the guardrails — and the margin recovered, the repricing lag fell, and the supplier loop closed itself.”

— Why these results compound

See it on your floor

Walk the platform against your stack — three steps, one thread.

The same three steps the case study above traces — how the agents come up, what they do on day one, and the guardrails each move crosses — are written up on the platform tour. Pricing is the next read.

See how it works on your floorSee the three tiers

Process: how it works → pricing