Product Strategy Case Study - Product Manager 2, Operations Excellence
Here's how I think Cult.fit's operations stack - from center audits to its existing computer-vision attendance pilot - could evolve into a real-time operational intelligence system across 700+ centers.
Cult.fit operates 700+ fitness centers across India, each functioning as a self-contained business unit responsible for revenue, member experience, and day-to-day operations. The Operations Excellence charter exists to make sure every one of those centers delivers a consistent, high-quality experience at scale - and to give central teams real-time visibility into what is actually happening on the ground, rather than learning about problems after a member has already been affected.
I'm proposing the Cult Operations Intelligence Platform (COIP): an internal product that closes the loop between what happens physically inside a center - a trainer running late, a treadmill breaking down, a wet floor going unmarked - and the people who need to act on it, in minutes rather than hours.
"Workflow automation systems that orchestrate actions across the organization before a single member notices a problem." - the operating philosophy this case study is built around.
My argument is simple: Cult.fit already has the ingredients. It has an existing, mature analytics backbone; it has already piloted computer vision at the center level for attendance; it already has a defined operations hierarchy that owns audits, facility upkeep, and member experience. What's missing is the connective tissue - the real-time event layer, the automated workflow engine, and the tiered command-center view - that would let the company move from measuring problems after the fact to resolving them before members feel them.
I'll walk through the current operating model, the pain points worth solving, the product vision, and a phased roadmap - starting from a single-cluster MVP that extends technology Cult.fit has already built, rather than proposing a platform from a blank page.
Each Cult.fit center operates like a small, accountable business unit rather than a cost center. A Center Manager (or Associate Center Manager) is judged on a blend of three things:
Membership sales, renewals, upsells, and personal training revenue generated in and around the center.
Timely opening and closing, inventory management of consumables and equipment, cleanliness, and audit compliance.
Attendance, engagement, complaint resolution, and satisfaction (NPS).
Above the center level, a Cluster Manager typically owns 7–10 centers and is responsible for standardizing procedures and ensuring audits are defect-free, along with hitting cluster-level sales and operations targets. Regional and Operations Managers, in turn, own P&L across a broader footprint - balancing revenue growth against operating cost, and keeping centers "Always New" through consistent maintenance and upkeep.
This nested, franchise-like accountability model is precisely why operational visibility matters so much: with hundreds of semi-autonomous units, the difference between a well-run and a poorly-run center is invisible to headquarters unless it is deliberately instrumented.
Six recurring workflows define the operational rhythm of a center:
| Workflow | Key Activities |
|---|---|
| Open Centre | Unlock, staff attendance, equipment inspection, cleaning verification, reception setup |
| Run Classes | Trainer check-in, member check-in/attendance, class delivery, feedback capture |
| Membership Operations | Walk-in → trial → conversion → renewal → upsell |
| Facility Operations | Equipment upkeep, inventory, utilities, cleaning, maintenance requests |
| Incident Management | Issue logged → owner assigned → resolved → verified → closed |
| Close Centre | End-of-day reporting, inventory reconciliation, security, cleaning |
Much of the coordination across this hierarchy still relies on manual mechanisms: checklists, physical or spreadsheet-based audits, and WhatsApp escalation chains between Center Managers, Cluster Managers, and Regional teams. Attendance historically used QR-code check-in - a mechanism the company itself has identified as fraud-prone, since a code can simply be shared between people - which led to a facial-recognition-based pilot as a more tamper-resistant alternative.
The direct implication: a meaningful share of "operational truth" - is the trainer actually present, is the equipment actually working, was the audit actually passed - still depends on self-reported or manually verified data, rather than continuously and automatically sensed data.
The Operations Excellence product serves an unusually wide stakeholder base - from a housekeeping associate with no smartphone habit at all, to a headquarters leader who wants a single number summarizing 700+ centers. Designing for all of them at once is the central product challenge.
| Stakeholder | What they need | Primary interaction |
|---|---|---|
| Trainer | Fast, frictionless attendance & schedule visibility in short gaps between sessions | Mobile app, tablet |
| Housekeeping / Maintenance | Simple task assignment and completion confirmation | Mobile / tablet checklist |
| Center Manager | Single view of center health; ability to triage issues immediately | Mobile + web dashboard |
| Cluster Manager | Cross-center comparison; audit compliance tracking across 7–10 centers | Web dashboard, alerts |
| Regional / Ops Manager | P&L-linked operational health; escalation visibility | Web dashboard, weekly reviews |
| HQ / Governance | Real-time, aggregated visibility across all 700+ centers | Command-center dashboard |
| Member | Indirect - benefits from a center that simply always works | None (invisible layer) |
Checklists, spreadsheet audits, and WhatsApp remain the connective tissue for a large share of day-to-day coordination. This works at small scale but creates inconsistent execution across hundreds of centers, and leaves no structured data trail for analysis.
Because escalation typically passes through several human layers - Center Manager to Cluster Manager to Regional Manager - headquarters often becomes aware of a problem only after it has already affected member experience, rather than as it is happening.
Trainer attendance historically relied on QR codes, which the company found could simply be shared between people to fake a check-in - a direct example of how self-reported data breaks down at scale, and the reason a facial-recognition pilot was built as a more reliable alternative.
Cluster Managers are explicitly accountable for keeping audits "defect-free" - implying quality is still verified periodically through audits rather than continuously through instrumentation. Problems like a broken machine or a missed cleaning task can go undetected between audit cycles.
Regional Operations Managers are tasked with keeping centers "Always New," but without continuous facility-health signals, this standard depends heavily on manager diligence and travel cadence rather than being systematically enforced.
Cult Operations Intelligence Platform (COIP)
Every operational event inside every Cult center should be measurable, predictable, and actionable in real time - resolved before a single member notices.
Rather than treating IoT, computer vision, workflow automation, and real-time dashboards as four separate initiatives, I'm framing them as one continuous loop:
IoT sensors, CCTV/vision-tech, app and check-in events capture what is physically happening
Analytics and lightweight AI models interpret those signals into operational meaning
A rules engine (and later, a recommendation layer) determines what action is warranted
A workflow engine assigns, notifies, and escalates automatically
Dashboards and alerts confirm resolution and feed the next audit cycle
Crucially, I'm not proposing a green-field platform. Cult.fit has already documented a mature analytics stack (a centralized warehouse and Metabase-based dashboards used by hundreds of employees daily, as of a 2020 engineering write-up) and has already piloted a working computer-vision product at the center level for attendance (as of a 2022 write-up). My job with COIP is to extend this foundation with a real-time, floor-level event layer - not to replace it.
COIP should be measured on operational leading indicators, not just lagging business metrics. Proposed KPI catalogue:
| KPI | What it measures | Owner |
|---|---|---|
| Center Health Score | Composite of equipment uptime, cleanliness, staffing, and audit status | Center Manager |
| Operational Latency | Time from issue occurrence to detection to resolution | Ops Excellence Product |
| Trainer Attendance Accuracy | % of attendance verified via automated (non-self-reported) means | Cluster Manager |
| Equipment Uptime % | Share of equipment operational vs. total inventory | Facility / Maintenance |
| Audit Defect Rate | Defects found per audit cycle, trending over time | Cluster Manager |
| Class Occupancy | Booked vs. attended vs. capacity per class | Center Manager |
| Member NPS / Complaint Rate | Experience quality as perceived by members | Regional Ops |
| Time-to-Resolution (Incidents) | Median and P90 time to close an incident end-to-end | Ops Excellence Product |
The goal of the IoT layer is to remove manual reporting for anything that can instead be sensed automatically. Because Cult.fit has already deployed tablet-based, on-device inference for attendance (a lightweight TFLite model on Android hardware already installed at pilot centers), the fastest-to-value IoT expansion is to extend this existing device footprint - rather than introducing an entirely new hardware layer - before investing in dedicated sensor hardware for equipment and occupancy.
| Signal source | Use case | Priority |
|---|---|---|
| Smart equipment / usage sensors | Detect equipment downtime and usage patterns without manual logging | High |
| Access control / RFID / NFC check-in | Verify member and staff presence automatically | High (partly live via facial check-in) |
| Occupancy sensors | Real-time class and floor occupancy without manual headcounts | Medium |
| Electricity / utility meters | Detect abnormal consumption suggesting equipment or facility issues | Medium |
| Temperature / environment sensors | Ensure comfort and safety thresholds are met | Low–Medium |
Cult.fit has already built and piloted an on-device facial recognition system for class attendance, designed to replace QR-based check-in. Registration captures facial embeddings (a 192-dimension vector via a MobileNet-based model, ~5.2MB and running on-device), and verification matches new captures against stored embeddings using a KNN/distance-based classifier.
This precedent matters to me strategically: it proves the organization can ship biometric CV at the center level, on affordable Android hardware, without heavy cloud inference costs. I want COIP's CV strategy to extend this same edge-first philosophy to operational (not just member-facing) use cases.
| Use case | Signal extracted | Complexity |
|---|---|---|
| Trainer presence detection | Is a trainer physically present for a scheduled class | Medium |
| Class / floor occupancy | Real member count vs. booked count | Medium |
| Housekeeping verification | Was a scheduled cleaning actually performed | Medium–High |
| Equipment utilization | Which machines are idle vs. in use, at what times | High |
| Safety / incident detection | Falls, crowding, or unsafe equipment use | High |
Workflow automation is the layer that turns a detected signal into a completed action without waiting for a human to notice, escalate, and manually assign it.
Each workflow should be designed with a visible owner, a time-bound SLA, and an automatic escalation path - replacing today's implicit, WhatsApp-driven escalation with something structured and measurable.
The Command Centre is the human-facing layer of COIP - a tiered dashboard giving each stakeholder the right altitude of visibility.
| Level | View | Primary use |
|---|---|---|
| HQ / Governance | All 700+ centers, aggregated health & alerts | Strategic oversight, exception spotting |
| Regional | Centers within region, P&L-linked ops health | Resource allocation, escalation review |
| Cluster | 7–10 centers, audit and SLA compliance | Coaching Center Managers, standardization |
| Center | Single-center live status | Daily triage and task management |
Importantly, I don't think Cult.fit needs to build this from scratch. As of a 2020 engineering write-up, the company already ran a centralized analytics stack - built on AWS Redshift and Postgres for near-real-time needs, ingesting over 1.5 billion events a month through 100+ ETL pipelines and 10+ sources, with Metabase serving as the shared dashboard layer for an 800+-user base (roughly 450 daily active users running 200,000+ queries a day), including Slack alerts and daily KPI emails.
I'd position COIP's Command Centre as a new real-time operational layer that plugs into this existing warehouse and dashboard investment - adding a lower-latency ingestion path for floor-level events (via a message broker such as Kafka or MQTT) alongside the existing batch/near-real-time pipelines - rather than a parallel data platform.
Solution design sketch - how the Observe / Understand / Decide / Act / Verify loop from Section 6 maps onto concrete architecture layers, and where each layer plugs into what Cult.fit has already built.
Rather than proposing a generic stack, I want the architecture to explicitly build on what Cult.fit's engineering team has documented investing in and validating in production: Redshift/Postgres for the warehouse, Metabase for dashboards and alerting, S3 (with Redshift Spectrum) as a data lake for cold storage, Hevo as an ETL layer for third-party sources, and lightweight on-device ML (TFLite-class models) for edge inference where latency and cost matter - as already piloted in the facial-attendance system.
Comparable operators at much larger scale validate the direction of this roadmap. Domino's built its store-operations visibility on a real-time streaming platform (piloted with Talend and Apache Kafka/Confluent across roughly 25 U.S. stores from 2022) so that franchise owners and store teams could see order volume and efficiency metrics as they happened, rather than after the fact - the same architectural bet this case study proposes for center-level events. Kripakar Krishnamurthy, Domino's Director of Enterprise Data Engineering, described the goal in terms worth borrowing directly: giving store managers "a smart watch for their in-store operations" - reliable, live information about what's happening right now, not a historical report.
Domino's has since layered AI on top of that real-time backbone, including a generative AI assistant (built with Microsoft's Azure OpenAI Service) designed to help store managers with inventory, ordering, and staff scheduling, plus ML-based delivery-route optimization and order-readiness prediction models reported to have improved from roughly 75% to 95% accuracy - evidence that once a real-time data layer exists, the AI opportunities on top of it compound quickly.
| Opportunity | Description |
|---|---|
| Predictive maintenance | Use equipment usage/failure history to flag machines likely to fail before they do |
| Anomaly detection in CV feeds | Flag unusual occupancy, safety, or presence patterns without hand-coded rules for every case |
| Natural-language summarization | Convert dashboard and incident data into a plain-English daily brief for HQ leadership |
| Smart escalation prioritization | Rank open incidents by likely member-experience impact, not just age |
| Root-cause clustering | Group recurring incidents across centers to surface systemic (vs. one-off) issues |
Given the fairness issue Cult.fit's own team already identified in its facial-recognition work, I think any new AI/CV capability under COIP needs an explicit accuracy-and-bias evaluation step before rollout, using a demographically representative internal dataset. Industry guidance on IoT/AI rollouts at scale is consistent on a second point too: anomaly-detection models should be introduced deliberately and tuned for low false-positive rates, since alert fatigue - not lack of data - is the most common reason frontline teams stop trusting an operational AI system.
Industry pattern across retail, QSR, and manufacturing IoT rollouts is consistent: start with one site (or one line, one store, one cluster), one clearly-painful problem, and a short pilot window - typically 6–8 weeks for a narrow pilot, up to 3–6 months for a fuller single-site pilot - before committing to a wider rollout. Domino's own real-time operations platform began as a single-purpose store-analytics use case before expanding into marketing and international scale. My MVP follows the same discipline: prove the operational-latency thesis on a single cluster and a single workflow before any large infrastructure investment.
I deliberately benchmarked the phase durations below against comparable real-world rollouts rather than assuming them: retail and manufacturing IoT programs typically run small pilots in 3–6 months and reach full multi-site scale only after 12–24 months, expanding in reusable "waves" that reuse the architecture proven in the pilot rather than redesigning it each time.
| Phase | Scope | Duration (indicative) |
|---|---|---|
| Phase 0 - Pilot | Single cluster; trainer-absence workflow; Center Health dashboard MVP | 0–3 months |
| Phase 1 - Regional Rollout | Expand workflow automation to 2–3 additional workflows (maintenance, cleaning); roll out to a full region | 3–7 months |
| Phase 2 - Vision-Tech Expansion | CCTV-based detection for occupancy and trainer presence in flagship centers | 7–12 months |
| Phase 3 - Full-Scale IoT + Command Centre | Sensor rollout across equipment; HQ-level real-time Command Centre live across all centers | 12–18 months |
| Phase 4 - Predictive Layer | Predictive maintenance, anomaly detection, smart escalation prioritization | 18+ months |
I'd have each wave deliberately reuse the prior phase's architecture (same event schema, same workflow engine, same dashboard layer) rather than re-platforming - this is the single biggest determinant of whether retail and manufacturing IoT rollouts stay on budget as they scale past the pilot site, per implementation guidance from comparable large-scale programs.
| Risk | Mitigation |
|---|---|
| CV model bias across skin tones / lighting (already observed internally) | Mandatory fairness evaluation on representative datasets before any new CV launch |
| Member and staff privacy concerns around CCTV-based monitoring | Clear consent and data-retention policy; use aggregate/derived signals, not stored raw video, wherever possible |
| Frontline change management - new tools add friction for trainers/housekeeping | Design for near-zero extra taps; pilot with heavy frontline feedback loops before wider rollout |
| Cost of sensor/camera hardware at 700+ center scale | Phase rollout by center tier/flagship status; prioritize software/workflow wins before hardware capex |
| Over-reliance on rules engine leading to alert fatigue | Start with a small number of high-confidence workflows; expand only as false-positive rates are proven low - flagged as the top reason frontline teams abandon operational AI/IoT systems in comparable retail and manufacturing rollouts |
At the platform level, I'd judge COIP's success against a small number of outcome metrics that tie directly back to the operational-latency thesis:
The domains differ, but the underlying product thinking transfers directly: both are about managing distributed, semi-autonomous human operations toward a consistent outcome, using data and workflow design rather than headcount.
| AECC (Education Counselling) | Cult.fit (Fitness Operations) |
|---|---|
| Counsellor | Trainer / Center Manager |
| Meeting → Application → Offer → Enrollment | Check-in → Class → Attendance → Renewal |
| Distributed counsellor performance visibility | Distributed center performance visibility |
| Conversion and retention optimization | Renewal and retention optimization |
| Manual follow-up workflows | Manual escalation via WhatsApp/audits |
This mapping - distributed operations, resource utilization, and conversion/retention optimization through structured product thinking rather than domain-specific fitness expertise - is the throughline I'd bring to this role: the domain is new to me, but the operating pattern isn't.