The visitor model records the variant in a boolean field per experiment, and conversion events such as a form submission, a signup or a purchase are tracked against the same visitor ID. A management command exports the pairs as CSV, and the product team analyses them in the tool they already use. No dashboard was built, and none was missed.
Services
Backend Engineering, Product
Industry
B2B Marketplace
Year
2022-2024
A/B tests inside a Django monolith.
The product team had argued about features for weeks: a new contact form, a simpler pricing display, an onboarding flow. Nobody could prove their version would convert better, and every change meant a full deploy, with a second one to take it back. What they wanted was easy to say: ship both versions, let the next 500 visitors decide, move on.
THE CHALLENGE
Every deploy was an all-or-nothing bet
The platform had no experimentation infrastructure. A change went to 100 % of visitors or to nobody, and a rollback was another release with another wait. There was no way to show a feature to 20 % of visitors and compare, so decisions were made on opinion and seniority. A dedicated experimentation platform was out of scope: no hosted service, no new infrastructure. The answer had to live inside the existing Django stack and be operated by the product team without a developer in the loop.
THE SOLUTION
Percentage flags tied to the visitor, not the session
I chose django-waffle over a hosted service like LaunchDarkly. A hosted service evaluates flags over the network on every request, needs an SDK in the process and bills per monthly active user; waffle evaluates in-process and stores flags in the same PostgreSQL database as everything else. Percentage flags are the core: the visitor-tracking middleware already gives every visitor a deterministic hash, and the flag's percentage is the threshold that hash is compared against. A context processor exposes the flag state to every template, and the same flag is checked in Python wherever the backend has to behave differently. A custom FlagAdmin makes the percentage editable in the list view, so a product manager moves a test from 5 % for internal QA to 10 %, 50 % and 100 %, or back to 0 %, in one save. The trade-off: waffle brings no analytics. The team took the switch and did the analysis elsewhere.
The admin customisation that puts the rollout percentage into the list view:
Python
class FlagAdmin(WaffleFlagAdmin):
list_display = [
'name',
'everyone',
'percent',
'superusers',
'staff',
'authenticated',
'note',
'created',
'modified',
]
list_editable = [
'everyone',
'percent',
'superusers',
'staff',
'authenticated',
'note',
]
list_filter = ['everyone', 'superusers', 'staff']Same visitor, same variant
Visitors land in A or B by their hash against the rollout percentage. Press "Visitor comes back" to re-send one of them, logged in or not: the bucket does not move. Then switch the flag off: everyone sees A, the hashes stay, and switching it back on restores every assignment.
show_new_contact_formonRulehash < 50 → B
v-0001hash 58A
v-0002hash 22B
v-0003hash 81A
v-0004hash 0B
v-0005hash 32B
v-0006hash 29B
v-0007hash 2B
v-0008hash 47B
v-0009hash 7B
v-000ahash 7B
v-000bhash 62A
v-000chash 92A
Each visitor is hashed once. The bucket follows the hash, not the session.
12Visitors
4Variant A
8Variant B
67%Actual B share
0Came back, same variant
One template conditional, no Python change:
HTML
{# base template, contact form A/B test #}
{% load waffle_tags %}
{% flag "show_new_contact_form" %}
{# Variant B: simplified form #}
{% include "contact/_form_v2.html" %}
{% else %}
{# Variant A: original form #}
{% include "contact/_form_v1.html" %}
{% endflag %}THE RESULT
Arguments got shorter and three features never got built
The team went from opinion to measurement. The new contact form won with a clear conversion lift, a simplified pricing display cut the bounce rate by 11 %, and three planned features were dropped before development because a small rollout showed no interest. An idea now takes about two hours to become a live experiment, and a bad variant is gone before most visitors have seen it. The flag system caused no incident in that time.
KEY METRICS
14A/B tests in 18 months
+23%Conversion lift, contact form
30sAverage rollback time
CLIENT FEEDBACK
"Features we used to argue about for weeks now go live as two versions, and the next 500 visitors settle it. Three features we were sure about showed no interest at all, so we never built them."
Product Manager
B2B marketplace, product team
FOR YOUR PROJECT
- When it applies
You run a monolith, product and engineering keep disagreeing about what to build, and a hosted experimentation platform is hard to justify for a dozen tests a year. In-process flags give you the switch without the subscription.
- What to check
Whether you have a visitor identity that survives login and logout. Without it, percentage flags bucket sessions instead of people and the data will not hold up. Decide who may flip a flag in production before the first test, not after.
- What it needs
django-waffle, a hook into your visitor middleware, a context processor and a small admin customisation: no new infrastructure. The discipline costs more than the code. Every flag needs a ticket, an end date and a removal, or the codebase fills with dead switches.
FAQ
TECHNOLOGY STACK
Django
Python
PostgreSQL
Manuel Kasbarian - CEO, SophistiXWe have enjoyed working with Daniel for 10 years now. We highly appreciate his fast response times around the clock and his all-round knowledge. Whether server configurations or programming, he always has the right solution.
Follow in the footsteps of Manuel and bring your vision to life.
Open to new ProjectsGet In Touch