Migrating a 274-page WordPress site to static HTML: engineering case study.

Engineering case study

Migrating a 274-page WordPress site to static HTML without losing an indexed URL

Opus Interactive runs colocation and private cloud out of three US facilities. Their marketing site was 38,867 WordPress files and 3.6 GB of plugin sprawl. We replaced it with a 274-page statically generated site, kept 227 of its 230 indexed URLs on their original paths, ported 207 redirect rules that were about to disappear, and cut over with roughly an hour of analytics gap and no measurable SEO loss.

Pages generated
274
Indexed URLs preserved
227 / 230
Redirect rules ported
207
Image payload cut
65%
Median TTFB
0.41 s
WordPress files removed from the serving path
38,867
The Opus Interactive homepage after the rebuild
The rebuilt homepage. Static HTML, no CMS, no database.
This is the engineering half. The commercial side of the same project, what the client was losing and what changed, is written up separately. Read the business case study →

The requirement: nothing that ranks may stop resolving

Opus Interactive has been selling infrastructure since 1996. Their buyers are CTOs and IT directors evaluating whether a provider will still be running their workloads in a decade. The site carried nearly thirty years of accumulated search equity: 95 blog posts, 135 pages, and a long tail of URLs from at least two previous site generations.

That made the brief unusual. The design work was the easy half. The hard requirement was this:

Nothing that currently ranks may stop resolving.

A conventional rebuild treats redirects as a launch-day chore. On a domain this old, redirects are the project. Everything below follows from taking that seriously.

Architecture: a build system, not a CMS

We did not replace WordPress with another CMS. The site is generated by a Python build system and deployed as flat files. There is no database, no application server in the request path, and no plugin surface to patch.

SOURCE _content/pagesextracted WP content, JSON _build/_*.htmlhead, header, tail, booking redirects-import.csv207 legacy rules build.py constantsSITE_BASE, CSS_VER GENERATE all.py 18 generators, 19 steps about · certs · policies datacenters · facility service · onepage · forms quote · location · bloggen industrieshub · build sitemap · notfound htaccess · seofiles 9,441 lines of Python one write_page() choke point OUTPUT 274 HTML pages+ sitemap, llms.txt .htaccess306 lines, generated kb-index.jsonAI search corpus SHIP deploy.py sha256 manifest diff, then tar|ssh verify every file additive, never deletes

Two design decisions did most of the work.

One choke point. Every page is written through a single write_page() function. Sitewide transforms (heading case normalisation, canonical stamping, cache-buster versioning) apply there rather than in eighteen places. Pages that have been hand-edited past the point where regeneration is safe carry a HAND-MAINTAINED marker and are skipped, with dedicated sweep functions to push shared changes into them anyway.

One constant per environment fact. SITE_BASE is the clearest example. The site previews in a subdirectory and lives at a domain root. Rather than maintain two builds, one constant drives the <base href> on all 274 pages, the RewriteBase in .htaccess, and whether the legacy redirect block is emitted at all.

# build.py
SITE_BASE = '/'          # preview: '/opus/'   live: '/'

# deploy.py refuses a mismatch in either direction
EXPECT_BASE = {'preview': '/opus/', 'live': '/'}
if want and base != want:
    sys.exit("REFUSING TO DEPLOY: build.py SITE_BASE is %r but the %r "
             "target needs %r." % (base, tname, want))

That guard exists because the failure is silent and total. A preview-shaped build pushed live would 301 every clean URL into a directory that does not exist on that host. A live-shaped build pushed to preview breaks every relative asset. Neither throws an error; both just serve a broken site.

Auditing every indexed URL before cutover

Before cutover we enumerated every URL the live WordPress site was publishing, from its own Yoast sitemaps, and resolved each one against the generated site.

== POSTS: 95 live URLs, 95 served by the new site, 0 MISSING

== PAGES: 135 live URLs, 132 served by the new site, 3 MISSING
   /opus-intearctive-hybrid-multicloud-solutions-news-and-resources  *** NO PAGE ***
   /infographics/                                                    *** NO PAGE ***
   /cloud-services-duplicate/                                        *** NO PAGE ***

All 95 post URLs and 132 of 135 page URLs resolve at byte-identical paths. The three exceptions are a typo'd slug, an empty archive and a duplicate, each given an explicit 301. Nothing that ranked was left to chance, and nothing was redirected that did not need to be.

The 207 redirects that were about to vanish

This is the part a sitemap diff cannot find, and the part most migrations get wrong.

Opus had a WordPress redirect plugin holding 207 rules built from their 404 logs over several years: an old /articles/*.html generation, category and tag URLs, press release paths, a pre-rebrand service tree. Those rules are PHP. The moment WordPress stops serving, all 207 stop with it.

They are also invisible to the obvious audit. Those URLs do not appear in the sitemap precisely because the plugin already redirects them. A sitemap diff reports the site as clean while 207 live redirects are minutes from deletion.

We ported them into Apache, generated from the client's own CSVs at build time so there is one source of truth:

# htaccess.py — emitted only when SITE_BASE == '/'
for src, tgt, code in legacy_redirects():
    pat = re.escape(src.lstrip('/'))
    a('  RewriteRule ^%s/?$ %s [R=%s,L,NC]' % (pat, tgt, code))

Validating the targets mattered as much as porting the rules. 66 of the 207 pointed at blog category archives that the rebuild had not produced, which would have turned working redirects into 301s to a 404. Rather than retarget them at a generic blog index and lose the topical signal, we generated 34 category archives at their original WordPress URLs, built from each post's full category assignments, marked noindex, follow to match exactly what Yoast had been emitting.

Verified after cutover

/category/vmware/       301 → /opus-blog/category/vmware/
/category/artificial-intelligence/  301 → /opus-blog/category/ai/
/it-services            301 → /managed-it-services/
/careers                301 → /about-opus/careers/
/bare-metal-cloud/      301 → /dedicated-iaas/

All 207 resolve single-hop onto pages that return 200. Zero redirect chains.

Why 207 rules were worth this much effort. The client-facing version covers how the redirect map was audited, graded and approved before anything was applied. Read the business case study →

Deploying to a host without rsync

The target server had no rsync binary and we were not going to install software on a client's production box to suit our tooling. So we built the part of rsync that mattered: diff by content hash, ship only what changed, verify what landed.

sha256, local504 staged files sha256, remoteone ssh call, same paths diffunchanged / changed / new tar | sshone stream, changed only verifyre-hash what landed

Every deploy ends with the server re-hashing what it received and comparing against the local digest. A deploy that cannot prove what landed is not a deploy, it is an upload.

== 539 files staged
== unchanged 537 | changed 2 | new 0
== uploaded 2 file(s)
== verified 2/2 by sha256

The verifier bug that only appeared with one file

The first end-to-end transport test reported TRANSPORT BROKEN. The file had actually arrived intact. The verifier was lying.

script = 'cd %s || exit 1; while IFS= read -r f; do '
         'if [ -f "$f" ]; then sha256sum "$f"; fi; done'
subprocess.run(ssh + [host, script], input='\n'.join(rels))   # ← no trailing newline

read returns non-zero at EOF on a line with no terminator, so the loop exits before the body runs on the final entry. With a multi-file list the last path was silently never hashed: it looked absent, was re-uploaded on every deploy, and the post-upload verification skipped it. With a single-file list, nothing was hashed at all, which is how a working upload reported as a failure.

One character fixed it. We found it only because we tested the transport with one file before trusting it with five hundred.

Designing the cutover so it could be undone

The site went live in one step, but the step was engineered to be reversible and to minimise the window where the domain served neither site cleanly.

Files ship in sorted order, which puts .htaccess first. Ship everything at once and Apache starts serving the new rules within a second, while the 500 pages those rules point at are still streaming in over the next two minutes. So six files that decide which site the domain serves are held back by default:

CUTOVER_FILES = ('.htaccess', 'robots.txt', 'llms.txt', 'favicon.ico',
                 'index.html', 'sitemap.xml')

index.html is on that list because of something we only knew by measuring. We probed the server's DirectoryIndex order by placing both an index.html and an index.php in a throwaway directory and requesting it:

GET /_diprobe/  →  HTML-WINS

HTML takes precedence. Dropping our index.html into the docroot would have switched the homepage instantly, with WordPress still otherwise live and no .htaccess change. Assuming the opposite would have produced a half-migrated homepage during what was supposed to be a silent staging step.

WordPress was not deleted. Its 38,867 files stayed on disk, with the front controller and admin paths denied in the new .htaccess, so rollback was two files and a few seconds rather than a 3.6 GB restore. The database was never touched.

Forms, abuse controls, and proving the rate limiter keys correctly

Six forms and a booking widget post to a single PHP endpoint that creates a Salesforce lead, emails the relevant inbox, and sends the visitor a branded confirmation. Because it sends a confirmation to whatever address is submitted, an unthrottled endpoint is a spam relay running on a freshly warmed sending domain.

Three layers, all failing open:

LayerBehaviourWhy
HoneypotSilent successField is hidden from humans, so a filled one is a bot with no human to disappoint
Time trap, 2s429, asks the visitor to send againA fast human is real. A silent drop means the visitor reads "your request is in" and nobody receives it
Rate limit, 5 per 10 min429 with a phone numberProtects the sending domain, not just the inbox

The rate limiter contained the project's highest-consequence assumption. The site sits behind a WAF, so every request reaches PHP from the proxy's edge address. Keying on REMOTE_ADDR would have put the entire internet in one bucket and locked out every visitor after five submissions, site-wide, silently.

We did not assume the header was right. We exhausted the limit through the proxy, then hit the origin directly and checked which bucket we landed in:

through-proxy #1..#4  422   (validation, counter incrementing)
through-proxy #5      429   ← limited
direct-to-origin      429   ← SAME bucket ⇒ keying on the real client IP

Same bucket means the limiter sees the visitor, not the edge. That is a two-minute test that de-risks a failure which would have been invisible until lead flow quietly stopped.

Performance: auditing by total bytes, not per page

BEFORE · WORDPRESS AFTER · STATIC HTML REMOVED 3,127 KB 985 KB 69% 82 requests 22 requests same page, same content Theme, page builder and plugin code The actual page: content, images, type Before, on WordPress 0 500 1,000 1,500 2,000 2,500 Images 2,476 KB Revolution Slider 202 KB WPBakery page builder 130 KB Textron theme 121 KB Page HTML 75 KB Fonts and other 71 KB Review slider plugin 45 KB Other plugins 8 KB After, rebuilt as static HTML 0 500 1,000 1,500 2,000 2,500 Images 812 KB Stylesheet 69 KB Fonts 65 KB Page HTML 39 KB JavaScript files none Transfer sizes as a browser receives them, gzip and brotli included. Measured from the archived page and the original files, which are still on the server. Third-party analytics and ad tags are excluded from both sides: unchanged by the rebuild.
Homepage weight, before and after. The bars show the site's own assets only; third-party analytics and ad tags sit outside this comparison because that container was carried across unchanged.

A static site removes the application server from the request path, which is most of the win. The rest was image delivery.

Third-party page audits flagged a 652 KB hero on the page being tested. Auditing by total bytes served rather than per page told a different story:

prefoot-servers739 → 283 KB × 268 pages
hero-cages651 → 213 KB
hero-healthcare557 → 231 KB
sol-colo416 → 124 KB
hero-banking362 → 94 KB

before   after (WebP q80)

The audit named a file on 4 pages. The largest actual cost was a pre-footer background on 268 pages, which no single-page report surfaces. Converting the set cut 5.4 MB to 1.9 MB, a 65% reduction, with no markup rewrite beyond repointing references.

Measured on the live siteValue
TTFB, interior pages0.41 s
Total response, interior pages0.61 s
Homepage HTML, uncompressed155 KB
Homepage HTML, gzip39.6 KB
Homepage image payload1.66 MB across 26 images
The industries hub page The colocation page
Generated pages. Every second-level page shares one hero component driven by per-page CSS custom properties for crop and scrim opacity.
What the speed gain was worth commercially. 39 scripts down to 1, and why a hosting company with a slow site was arguing against itself. Read the business case study →

Two accessibility defects a visual review cannot catch

An agent-accessibility audit surfaced two defects that a visual review would not catch.

The first was a genuine keyboard trap. A decorative globe container carried aria-hidden="true" to hide the photo, but three interactive location buttons lived inside it. Keyboard users could tab onto controls that screen readers had been told did not exist. The photo already had alt="", so the correct fix was removing aria-hidden entirely rather than making the buttons unfocusable.

The second was invalid ARIA on six animated counters. A bare <span> has an implicit role of generic, which does not permit an accessible name, so the aria-label carrying the final value was being ignored. Adding role="img" makes the name valid and means assistive technology announces "200+" rather than whatever digit the odometer is mid-animation on.

- <span class="accent counter" data-count="200" aria-label="200+">
+ <span class="accent counter" data-count="200" role="img" aria-label="200+">

Five defects, and the checks that caught them

Every one of these was found before launch, by a check that produced a number we could compare against an expectation. None was found by looking at the site. That is the part worth copying, so the detection method is listed beside each defect.

IssueRoot causeCaught by
Every deep missing URL returned 500, not 404MultiViews resolves /contact to contact.html and passes the rest as PATH_INFO, so REQUEST_FILENAME.html exists and the clean-URL rule recurses until Apache abortsPost-deploy URL sweep
Sentence-case pass lowercased proper nounsA sitewide heading transform with no proper-noun protection: "Shannon hulbert" on the leadership page, "manassas" in six service headingsClient review, then a sitewide sweep
Cache-buster grew without bound[a-z0-9]+ in the version-rewrite regex did not match the hyphen in a new version string, so each build appended again: ?v=20260928-live-live-live-liveIn-browser stylesheet check
A content fix silently revertedEdited generated HTML instead of the source JSON; the next build overwrote it and the file then matched the server, so it never appeared in the deploy diffChanged-file count being one lower than expected
macOS tar smuggled AppleDouble filesExtended attributes produce ._name siblings that Apache will serveListing the probe directory after the first transport test

The deploy manifest alone caught two of them, purely because the changed-file count did not match what we had predicted. A build pipeline that reports counts you can predict is worth more than one that reports success.

Stack

Before and after

BeforeAfter
PlatformWordPress, 38,867 files, 3.6 GB274 static pages
Request pathPHP + MySQL + plugin stackFlat files
Attack surfacePlugins, themes, admin, XML-RPCTwo PHP endpoints
Indexed URLs230227 identical, 3 redirected
Legacy redirects207, in a plugin207, in Apache
Image payload5.4 MB1.9 MB
AnalyticsGA4 + GTMIdentical, both properties preserved

Call tracking, session recording and both ad pixels survived the migration untouched, because the tag container was carried across intact rather than rebuilt. Verifying what was actually inside it before touching it is the only reason that held.

The rest of the project. Content extraction, the booking system, the icon set and the visual work are covered in the business case study. Read the business case study →

Working on something with these constraints?

Migrations where the URLs matter, build systems for sites too large to hand-maintain, or integrations that need to survive a platform change. If that is the shape of your problem, we should talk.

Start a conversation