Case Study - Abercrombie & Fitch

Abercrombie & Fitch wakes up to a finished smoke suite across six regions

A conversation between
Sneha Sivakumar
Sneha Sivakumar
CEO of Spur
Lauren Morr
Senior Vice President, Digital, Abercrombie & Fitch Co.

COMPANY

Global specialty retailer — A&F, abercrombie kids, Hollister, Gilly Hicks

INDUSTRY

Global Specialty Apparel Retail

COMPANY SIZE

10,000+

FOUNDED

1892

4AM

Daily smoke execution,

fully automated, replacing a manual overnight run.

23

Production defects caught

across 10 daily smoke suites in the first 90 days.

6

Regions watched daily,

plus 8 franchise partner sites under hyper-care.

The Problem

Two brands, six regions, a deploy every day, and a smoke suite run by hand

Abercrombie & Fitch Co. is a global specialty retailer operating Abercrombie & Fitch, abercrombie kids, Hollister and Gilly Hicks across North America, EMEA and APAC. Code ships every day, which means the site has to be verified every single morning across six regional storefronts, eight franchise partner sites and more than 40 supported languages.

The smoke suite that gated those deploys was run by hand, overnight, and validation sat with the QE team alone. Franchise and localized sites were checked ad hoc. Analytics faults could sit undetected for a month. And on mobile, iOS regression was one week of fully manual testing, against 14,000–15,000 unit tests per App Store update and monthly releases.

What Changed

A narrow pilot, expanded into five teams

Abercrombie didn't start with the whole estate. The pilot covered two validation types, stage prop validations and analytics events, chosen because they were measurable and low-risk. Agent runtime on the same test fell from 17 minutes to 8 minutes.

Onboarding converted what already existed rather than rewriting it: pre-prod validation steps that lived in a spreadsheet, and Playwright scripts, came across as end-to-end Spur flows. The daily pipeline now makes all P0 tests compulsory before deploy, Abercrombie and Hollister, desktop and mobile, US and EMEA, with results posting to Teams and failures triaged before the release gate.

The smoke suite runs itself at 4AM, fully automated, finished before anyone opens a laptop. From there, validation expanded to five teams: QE, Martech, engineering, mobile QE and franchise. A 3-day mobile POC put the one-week regression cycle on a path to roughly two hours, and one engineer automated four of eight franchise sites at fifteen test cases each, displacing about five manual hours per franchise per week.

What it Unlocked

23 production defects that would have reached customers

In the first 90 days, the ten daily smoke suites caught 23 unique production defects. The specifics show why visual, agent-driven testing matters at this scale: a cart sync failure where remove-from-bag silently failed, traced through network logs; the Qatar storefront serving imperial units instead of metric; a Mexico PDP rendering a 446cm model height; a German translation audit finding seven pages shipping an untranslated header and nav; a Quick View modal failing intermittently from a gap in the EU smoke suite; and Tealium order_confirmation cert failures in EMEA.

Analytics faults that could once sit undetected for a month are now caught the same day. These are exactly the faults that scripted selectors pass over, the page loads, the button exists, and the content is still wrong.

"I got the login and just tried it. Without any training I had a comprehensive report in an hour, it called out every functional difference between the two pages, including translated content I'd never have caught by myself."

The smoke suite that used to run on human hours now runs on a schedule, and success is measured the way leadership wants it measured: defects caught, coverage analyzed continuously, manual hours displaced, and a domain-by-domain rollout plan. The regression marathon didn't get faster; it stopped being a human job at all.

4AM

Daily smoke execution, fully automated

23

Production defects caught in 90 days

6

Regions watched daily, plus 8 franchise sites

Critical e-commerce flows across 30+ regions
Every regional price, discount rule, and product variant automatically tested before your sale goes live, no manual spot-checking required.
Hundreds of partner landing pages
Ensuring that every audience coming from podcasts, newsletters, and other partnerships lands on a page that is on brand and error free.
Staging and production environments
Running tests in staging for high confidence before launch, then validating again on production as a final safety net.

Key Insights

Two brands, six regions, a deploy every day. The smoke suite that gates those deploys used to be run by hand, overnight. It now runs itself — finished before anyone opens a laptop — and caught 23 production defects in the first 90 days.

CUSTOMER STORIES

More teams, same results.

Two brands, six regions, a deploy every day
Abercrombie & Fitch wakes up to a finished smoke suite across six regions
Six months of Playwright, beaten in a single day
Every release and every drop, validated
Launch QA done before the 10AM review
Every product page in a launch, checked before the team logs on
OneSafe Customer Card
QA time per release dropped from four days to one
How OneSafe Cut Release QA From Four Days to One
Launch-day QA went from three or four people running manual regression to one engineer
How Our Place Automated 80% of Its QA Coverage and Freed Its Team’s Time Up
UncommonGoods cut QA time in half with AI-driven testing
How Uncommon Goods stopped spending 50% of their QA time on Selenium
From Manual QA Bottlenecks to Fast, Reliable Releases with Spur
How Wondr Health enabled an entire team to work on more interesting problems
Scaling shoppable UGC QA across dozens of brands by adding a single URL to a shared Spur scenario table
How Hue QAs shoppable widgets across 20+ merchants without rebuilding anything.
Regression Done by Noon, Every Release
The regression marathon that ran 9am to midnight, every two weeks