Shipping an AI Photo Editor as a WordPress Plugin: 22 Versions of Hard Lessons

· 11 min read · case-studies wordpressphpfal-aidockerai-imagejavascript

A few weeks ago I wrote about the proof of concept: one index.html, no backend, two furniture photos in, a staged room scene out. It answered the question it was built to answer — can you keep the real product and only generate the room around it? — and it left one obvious question open.

That PoC ended with the words “before committing to the full WordPress plugin.” This post is about the plugin. It went from version 0.0.1 to 0.0.22 across 58 commits, and almost none of that work was the AI part.

Why it matters (for a business owner, not a developer)

The PoC’s limitation was that it lived in a single HTML file with the API key in the source. That’s fine for a demo you email a client. It is not fine for something a furniture shop puts on their own website, where:

Every one of those five points is a product requirement, and every one of them cost more code than the AI call itself.

Architecture: the boring parts that make it work

The plugin is a normal WordPress plugin — PHP prefix hre_, namespace HRE, custom tables created with dbDelta() on activation, an autoloader, no Composer runtime dependency. The interesting decisions are elsewhere.

Generation is asynchronous, always

fal.ai jobs take 15–20 seconds. PHP-FPM request timeouts, max_execution_time, and impatient visitors make synchronous generation a losing bet. So the flow is a queue-and-poll:

visitor uploads  →  POST /uploads         →  normalize → start vision analysis
visitor picks    →  GET  /prompts/{cat}   →  per-style prompts
visitor confirms →  POST /jobs            →  fal queue submit (returns uuid)
browser polls    →  GET  /jobs/{uuid}     →  status: queued | processing | done

The visitor never waits on a provider. The job row carries the fal request id, and a poll either finds a finished result or shows genuine progress. The front end is a small state machine over a session cookie, so a mid-wizard refresh restores where the visitor was.

Anonymous visitors have no accounts, which makes authorization easy to get wrong. The rule I settled on: every read resolves through the caller’s own session hash, never through an ID accepted from the client.

That’s one decision that removes an entire class of “I changed the id in the URL and saw someone else’s photo” bugs.

The watermark is a hard requirement, so a fallback must never be silent

This is the bug that taught me the most. The watermark step does a GD imagecopyresampled composite. If it ever fails, the tempting “resilient” move is:

// DON'T do this
if ( ! $watermarked ) {
    rename( $stage, $final );   // ship the un-watermarked file
}

I did exactly that. It’s silent, it has no log line, and it means the plugin happily serves un-watermarked results while the admin screen says “watermark: on.” The owner found out weeks later by noticing a photo without a watermark.

The fix was three parts:

  1. Log the root cause, not the nearest symptom. apply() returned hre_no_watermark_source — a downstream error. The real cause was that the watermark source resolution chain (site_logoget_theme_mod('custom_logo')) had no logo to resolve. I added a block_reason() diagnostic that re-walks the chain and returns a specific message per failure branch.
  2. Surface it in the admin. A persistent notice-warning says the watermark is enabled but won’t be applied, plus the concrete reason. No log-reading required.
  3. Make the failure visible to me in testing. A standalone probe composites a known watermark onto a known base and samples a pixel inside the expected watermark rectangle — proving the transform path itself works, which separates “the feature is broken” from “the feature is unconfigured.”

Rule I now follow: when a best-effort branch masks a non-negotiable transform, it must log the concrete reason and warn in the admin UI. A silent fallback that degrades a required feature is not resilience — it’s a bug with good manners.

Lead capture: off by default, asynchronous when on

Lead capture is a master toggle. Off, the plugin collects nothing and the wizard behaves exactly as before. On, the visitor’s configured fields are required before generation, stored in a custom table, and forwarded to a webhook.

Two decisions worth copying:

wp_schedule_single_event( time() + 5, 'hre_lead_webhook', array( $lead_id ) );

Delivery retries three times with linear backoff, then marks the record failed. The admin list shows a status badge (pending / retrying / delivered / failed) plus the attempt count, and — this matters — the plugin’s data retention rule keeps delivered leads but purges undelivered ones on the normal schedule, because an undelivered lead is just PII sitting in your database.

The provider kill switch has to be enforced server-side

The admin can disable the image provider and set a message that visitors see instead of the generate button. Greying out the checkbox is UX; the enforcement is a check at the top of the job-creation endpoint:

if ( ! (bool) Settings::get( 'fal_provided' ) ) {
    return $this->error_response( 'hre_provider_disabled', $msg, 503 );
}

Same for the lead gate. The browser gate is UX. The server gate is the product.

The five bugs that only showed up in the real install

This is the part no article about “building an AI plugin” ever includes, and it’s where most of the 58 commits went.

1. Admin REST routes registered inside is_admin() silently 404

Registering admin routes only when is_admin() is true looks correct and is completely broken: REST requests report is_admin() === false, so the routes are never registered and every call returns 404 "No route was found matching the URL and request method" — which reads exactly like a URL typo, not a registration bug. Register at plugin boot; the permission_callback (capability + nonce) is what gates access.

2. rest_url() has no trailing slash

The one that cost me the most recent afternoon. rest_url() returns https://example.com/wp-json/my-plugin/v1/ — the trailing slash is on the namespace, and the next path segment must start with its own slash. Concatenating without one:

const rest = window.hreAdmin.rest;      // ".../wp-json/hoelee-ai-photo-remix/v1"
fetch( rest + 'admin/photos/' + id )    // ❌ ".../v1admin/photos/123"
fetch( rest + '/admin/photos/' + id )   // ✅

.../v1admin/photos/123 is a 404 with no browser console error, so the catch block fires and the UI shows the generic “Something went wrong.” It had quietly broken four endpoints — photo pagination, photo delete, the webhook test button, and lead resend. The fix is one character per call site; the diagnosis took far longer, because the failure mode is indistinguishable from a server error. Prove it with two curls: the correct join returns 403 (route exists, nonce missing), the broken one returns 404.

3. A WordPress admin settings form can wipe a different tab’s settings

Each admin tab posts only its own fields, but the save handler runs the whole settings map. A checkbox that isn’t submitted looks identical to a checkbox that was unchecked, so saving the “limits” tab silently wrote false over the “share by default” option, and saving that one cleared the entire lead-capture configuration. The fix belongs in the sanitizer, not the forms — distinguish “absent from this submission” from “present and empty”:

// present in this tab  → use the new value (a cleared field clears)
// absent from this tab → keep the current value
if ( array_key_exists( $key, $raw ) ) {
    $clean[ $key ] = sanitize( $raw[ $key ] );
}

I added a regression test for exactly this (seed config → save a different tab’s subset → assert the seed survived). It’s the kind of bug that only exists once you have two tabs, which is why it shipped.

4. A binary endpoint must not go through a JSON fetch helper

The result download is a JPEG stream. The shared api() helper did res.json() on every response, which threw on image bytes — and a swallowing .catch(function(){}) ate the error, so the symptom was “the result never shows up” with nothing in the console. Blob endpoints need their own raw fetch helper that checks res.ok, parses JSON only on error, and returns res.blob() on success.

5. A size limit that fires before the resize defeats the resize

The plugin downscales uploads to a working size (1920×1080 bounds). I had set the decompression-bomb guard too low — 1 MB / 1 MP — which meant every legitimate phone photo was rejected before it could be resized. Visitors saw the preview appear and instantly de-select, reported as “the resize feature is broken.”

The real lesson is UX, not numbers: the client must reject an over-limit file before showing the local preview. Flash-then-disappear reads as a bug even when the message underneath is correct. The caps now sit at 15 MB / 50 MP, high enough that normal photos always pass and get normalized.

What I’d do differently

The result

A production-shaped WordPress plugin: 0.0.22 across 22 releases, with 10 custom database tables, an async fal.ai queue, server-side key handling, dynamic lead capture with a retrying webhook, a role-aware photo archive, and everything configurable from a 10-tab admin screen without touching code.

It is the difference between a demo and a product, and — the number I actually care about — 58 commits, of which roughly four were about the AI model. The rest was the unglamorous work that decides whether a client can run the thing without me: validation, retention, permissions, failure visibility, and an admin screen that tells the truth.


Want this for your business?

I build WordPress plugins, AI image pipelines, and self-hosted infrastructure for Malaysian SMEs — and I ship them as products you can actually run, not demos you have to babysit. If you want AI product photography, a lead-capture pipeline, or a custom plugin for your business, I’d love to talk:

Lee Teong Hoe

Full-stack developer & DevOps engineer. I build web apps, self-host infrastructure, and automate things — this blog is my living portfolio.