Product Strategy Case Study - Product Manager 2, Operations Excellence

Designing the Cult Operations Intelligence Platform Closing the loop between what happens on the gym floor and who needs to act on it

Here's how I think Cult.fit's operations stack - from center audits to its existing computer-vision attendance pilot - could evolve into a real-time operational intelligence system across 700+ centers.

Prepared by Prem Sameer
sameer.prem1405@gmail.com
Last reviewed 11 July 2026
A note on how I built this. This is an independent case study I put together for the interview process - not an internal Cult.fit document, and nothing here relies on confidential information. Every specific claim I make about Cult.fit's existing systems (the facial-recognition attendance pilot, the analytics stack) comes from Cult.fit's own public engineering blog, and I've cited it by name and date throughout, because those posts are the last publicly documented state of those systems and things may well have moved on since. Where I lean on outside precedent (Domino's real-time operations platform, for instance), I've cited that too. My goal was to reason like a PM would on day one - from evidence, not assumption - rather than to claim insider knowledge of Cult.fit's current architecture.

01Executive Summary

Cult.fit operates 700+ fitness centers across India, each functioning as a self-contained business unit responsible for revenue, member experience, and day-to-day operations. The Operations Excellence charter exists to make sure every one of those centers delivers a consistent, high-quality experience at scale - and to give central teams real-time visibility into what is actually happening on the ground, rather than learning about problems after a member has already been affected.

I'm proposing the Cult Operations Intelligence Platform (COIP): an internal product that closes the loop between what happens physically inside a center - a trainer running late, a treadmill breaking down, a wet floor going unmarked - and the people who need to act on it, in minutes rather than hours.

"Workflow automation systems that orchestrate actions across the organization before a single member notices a problem." - the operating philosophy this case study is built around.

My argument is simple: Cult.fit already has the ingredients. It has an existing, mature analytics backbone; it has already piloted computer vision at the center level for attendance; it already has a defined operations hierarchy that owns audits, facility upkeep, and member experience. What's missing is the connective tissue - the real-time event layer, the automated workflow engine, and the tiered command-center view - that would let the company move from measuring problems after the fact to resolving them before members feel them.

I'll walk through the current operating model, the pain points worth solving, the product vision, and a phased roadmap - starting from a single-cluster MVP that extends technology Cult.fit has already built, rather than proposing a platform from a blank page.

02Business Model

Each Cult.fit center operates like a small, accountable business unit rather than a cost center. A Center Manager (or Associate Center Manager) is judged on a blend of three things:

Revenue performance

Membership sales, renewals, upsells, and personal training revenue generated in and around the center.

Operational excellence

Timely opening and closing, inventory management of consumables and equipment, cleanliness, and audit compliance.

Member experience

Attendance, engagement, complaint resolution, and satisfaction (NPS).

Above the center level, a Cluster Manager typically owns 7–10 centers and is responsible for standardizing procedures and ensuring audits are defect-free, along with hitting cluster-level sales and operations targets. Regional and Operations Managers, in turn, own P&L across a broader footprint - balancing revenue growth against operating cost, and keeping centers "Always New" through consistent maintenance and upkeep.

This nested, franchise-like accountability model is precisely why operational visibility matters so much: with hundreds of semi-autonomous units, the difference between a well-run and a poorly-run center is invisible to headquarters unless it is deliberately instrumented.

03Current Operating Model

3.1 Organizational Hierarchy

  • Headquarters / Central Operations & Product Leadership
  • Regional Manager - owns P&L across a region
  • Cluster Manager - owns 7–10 centers, standardizes SOPs, owns audit compliance
  • Center Manager / Associate Center Manager - owns a single center end-to-end
  • Trainers, Reception, Housekeeping, Maintenance - frontline execution

3.2 Core Operational Workflows

Six recurring workflows define the operational rhythm of a center:

WorkflowKey Activities
Open CentreUnlock, staff attendance, equipment inspection, cleaning verification, reception setup
Run ClassesTrainer check-in, member check-in/attendance, class delivery, feedback capture
Membership OperationsWalk-in → trial → conversion → renewal → upsell
Facility OperationsEquipment upkeep, inventory, utilities, cleaning, maintenance requests
Incident ManagementIssue logged → owner assigned → resolved → verified → closed
Close CentreEnd-of-day reporting, inventory reconciliation, security, cleaning

3.3 How Work Actually Gets Coordinated Today

Much of the coordination across this hierarchy still relies on manual mechanisms: checklists, physical or spreadsheet-based audits, and WhatsApp escalation chains between Center Managers, Cluster Managers, and Regional teams. Attendance historically used QR-code check-in - a mechanism the company itself has identified as fraud-prone, since a code can simply be shared between people - which led to a facial-recognition-based pilot as a more tamper-resistant alternative.

Source check: Cult.fit's engineering team documented this shift directly: "The QR based attendance is prone to fraud as QR codes can easily be shared with others." - "Minimizing Attendance Fraud Through Facial Recognition at Cult Centers," blog.cult.fit, May 2022.

The direct implication: a meaningful share of "operational truth" - is the trainer actually present, is the equipment actually working, was the audit actually passed - still depends on self-reported or manually verified data, rather than continuously and automatically sensed data.

04Stakeholder Analysis

The Operations Excellence product serves an unusually wide stakeholder base - from a housekeeping associate with no smartphone habit at all, to a headquarters leader who wants a single number summarizing 700+ centers. Designing for all of them at once is the central product challenge.

StakeholderWhat they needPrimary interaction
TrainerFast, frictionless attendance & schedule visibility in short gaps between sessionsMobile app, tablet
Housekeeping / MaintenanceSimple task assignment and completion confirmationMobile / tablet checklist
Center ManagerSingle view of center health; ability to triage issues immediatelyMobile + web dashboard
Cluster ManagerCross-center comparison; audit compliance tracking across 7–10 centersWeb dashboard, alerts
Regional / Ops ManagerP&L-linked operational health; escalation visibilityWeb dashboard, weekly reviews
HQ / GovernanceReal-time, aggregated visibility across all 700+ centersCommand-center dashboard
MemberIndirect - benefits from a center that simply always worksNone (invisible layer)

05Operational Pain Points

5.1 Fragmented, Manual Tooling

Checklists, spreadsheet audits, and WhatsApp remain the connective tissue for a large share of day-to-day coordination. This works at small scale but creates inconsistent execution across hundreds of centers, and leaves no structured data trail for analysis.

5.2 Delayed Visibility at Headquarters

Because escalation typically passes through several human layers - Center Manager to Cluster Manager to Regional Manager - headquarters often becomes aware of a problem only after it has already affected member experience, rather than as it is happening.

5.3 Manual, Fraud-Prone Verification

Trainer attendance historically relied on QR codes, which the company found could simply be shared between people to fake a check-in - a direct example of how self-reported data breaks down at scale, and the reason a facial-recognition pilot was built as a more reliable alternative.

5.4 Audit-Driven Rather Than Continuously-Sensed Quality

Cluster Managers are explicitly accountable for keeping audits "defect-free" - implying quality is still verified periodically through audits rather than continuously through instrumentation. Problems like a broken machine or a missed cleaning task can go undetected between audit cycles.

5.5 Inconsistent "Always New" Standard

Regional Operations Managers are tasked with keeping centers "Always New," but without continuous facility-health signals, this standard depends heavily on manager diligence and travel cadence rather than being systematically enforced.

06Product Vision

Cult Operations Intelligence Platform (COIP)

Every operational event inside every Cult center should be measurable, predictable, and actionable in real time - resolved before a single member notices.

Rather than treating IoT, computer vision, workflow automation, and real-time dashboards as four separate initiatives, I'm framing them as one continuous loop:

Observe

IoT sensors, CCTV/vision-tech, app and check-in events capture what is physically happening

Understand

Analytics and lightweight AI models interpret those signals into operational meaning

Decide

A rules engine (and later, a recommendation layer) determines what action is warranted

Act

A workflow engine assigns, notifies, and escalates automatically

Verify

Dashboards and alerts confirm resolution and feed the next audit cycle

Crucially, I'm not proposing a green-field platform. Cult.fit has already documented a mature analytics stack (a centralized warehouse and Metabase-based dashboards used by hundreds of employees daily, as of a 2020 engineering write-up) and has already piloted a working computer-vision product at the center level for attendance (as of a 2022 write-up). My job with COIP is to extend this foundation with a real-time, floor-level event layer - not to replace it.

07Key Performance Indicators

COIP should be measured on operational leading indicators, not just lagging business metrics. Proposed KPI catalogue:

KPIWhat it measuresOwner
Center Health ScoreComposite of equipment uptime, cleanliness, staffing, and audit statusCenter Manager
Operational LatencyTime from issue occurrence to detection to resolutionOps Excellence Product
Trainer Attendance Accuracy% of attendance verified via automated (non-self-reported) meansCluster Manager
Equipment Uptime %Share of equipment operational vs. total inventoryFacility / Maintenance
Audit Defect RateDefects found per audit cycle, trending over timeCluster Manager
Class OccupancyBooked vs. attended vs. capacity per classCenter Manager
Member NPS / Complaint RateExperience quality as perceived by membersRegional Ops
Time-to-Resolution (Incidents)Median and P90 time to close an incident end-to-endOps Excellence Product

08IoT Strategy

The goal of the IoT layer is to remove manual reporting for anything that can instead be sensed automatically. Because Cult.fit has already deployed tablet-based, on-device inference for attendance (a lightweight TFLite model on Android hardware already installed at pilot centers), the fastest-to-value IoT expansion is to extend this existing device footprint - rather than introducing an entirely new hardware layer - before investing in dedicated sensor hardware for equipment and occupancy.

Signal sourceUse casePriority
Smart equipment / usage sensorsDetect equipment downtime and usage patterns without manual loggingHigh
Access control / RFID / NFC check-inVerify member and staff presence automaticallyHigh (partly live via facial check-in)
Occupancy sensorsReal-time class and floor occupancy without manual headcountsMedium
Electricity / utility metersDetect abnormal consumption suggesting equipment or facility issuesMedium
Temperature / environment sensorsEnsure comfort and safety thresholds are metLow–Medium

09Computer Vision Strategy

9.1 Existing Foundation

Cult.fit has already built and piloted an on-device facial recognition system for class attendance, designed to replace QR-based check-in. Registration captures facial embeddings (a 192-dimension vector via a MobileNet-based model, ~5.2MB and running on-device), and verification matches new captures against stored embeddings using a KNN/distance-based classifier.

What the pilot actually reported (worth being precise about): per Cult.fit's own engineering post, the pilot ran live at a single Bengaluru center (Cult HSR 19th Main) with 120 trainers and 6 Center Managers onboarded for registration, and - notably - only 9 customers onboarded at the time of writing. Reported metrics were 99.7% accuracy, a ~15s mean registration time, and a 2.25s mean time to mark attendance. This is a real, working system - but the last public data point is a single-center pilot from May 2022. I'm treating it as a proven proof-of-concept to extend, not as evidence that facial attendance already runs network-wide today.

Source: "Minimizing Attendance Fraud Through Facial Recognition at Cult Centers," blog.cult.fit, May 4, 2022.

This precedent matters to me strategically: it proves the organization can ship biometric CV at the center level, on affordable Android hardware, without heavy cloud inference costs. I want COIP's CV strategy to extend this same edge-first philosophy to operational (not just member-facing) use cases.

9.2 Proposed CV Use Cases (CCTV / Vision-Tech)

Use caseSignal extractedComplexity
Trainer presence detectionIs a trainer physically present for a scheduled classMedium
Class / floor occupancyReal member count vs. booked countMedium
Housekeeping verificationWas a scheduled cleaning actually performedMedium–High
Equipment utilizationWhich machines are idle vs. in use, at what timesHigh
Safety / incident detectionFalls, crowding, or unsafe equipment useHigh

9.3 A Known Risk Worth Naming Early

Cult.fit's own engineering team has publicly acknowledged that its facial recognition feature-generation model showed a bias towards lighter-complexion faces in testing, and flagged retraining on a more representative dataset as future work. Any CV expansion under COIP should treat fairness testing across skin tones and lighting conditions as a launch requirement, not an afterthought - this is a real, previously identified risk, not a hypothetical one. (Same source as above.)

10Workflow Automation

Workflow automation is the layer that turns a detected signal into a completed action without waiting for a human to notice, escalate, and manually assign it.

Example: Trainer Absence

  • Signal: scheduled trainer does not check in within a defined grace window
  • Automated action: task created and assigned to Center Manager
  • Parallel action: affected members notified of delay or substitute trainer
  • Escalation: if unresolved within X minutes, automatically escalated to Cluster Manager

Example: Equipment Failure

  • Signal: usage sensor or manual flag indicates equipment down
  • Automated action: maintenance ticket created, technician notified, equipment marked unavailable in booking flows

Example: Missed Cleaning Task

  • Signal: scheduled cleaning task not confirmed (via CV or checklist app) within window
  • Automated action: housekeeping lead notified; task escalated if repeated at same center

Each workflow should be designed with a visible owner, a time-bound SLA, and an automatic escalation path - replacing today's implicit, WhatsApp-driven escalation with something structured and measurable.

11Operations Command Centre

The Command Centre is the human-facing layer of COIP - a tiered dashboard giving each stakeholder the right altitude of visibility.

LevelViewPrimary use
HQ / GovernanceAll 700+ centers, aggregated health & alertsStrategic oversight, exception spotting
RegionalCenters within region, P&L-linked ops healthResource allocation, escalation review
Cluster7–10 centers, audit and SLA complianceCoaching Center Managers, standardization
CenterSingle-center live statusDaily triage and task management

Importantly, I don't think Cult.fit needs to build this from scratch. As of a 2020 engineering write-up, the company already ran a centralized analytics stack - built on AWS Redshift and Postgres for near-real-time needs, ingesting over 1.5 billion events a month through 100+ ETL pipelines and 10+ sources, with Metabase serving as the shared dashboard layer for an 800+-user base (roughly 450 daily active users running 200,000+ queries a day), including Slack alerts and daily KPI emails.

Reading this number correctly: that snapshot is from January 2020, when Cult.fit operated roughly 250 centers - a third of today's 700+. The specific volumes have almost certainly grown well past 1.5B events/month since; what I'd take from this source is the pattern (centralized warehouse + Metabase + Hevo ETL, batch and 15-minute near-real-time tiers), not the exact 2020 figures.

Source: "cure.fit's Analytics Stack," blog.cult.fit, January 30, 2020.

I'd position COIP's Command Centre as a new real-time operational layer that plugs into this existing warehouse and dashboard investment - adding a lower-latency ingestion path for floor-level events (via a message broker such as Kafka or MQTT) alongside the existing batch/near-real-time pipelines - rather than a parallel data platform.

12System Architecture

Solution design sketch - how the Observe / Understand / Decide / Act / Verify loop from Section 6 maps onto concrete architecture layers, and where each layer plugs into what Cult.fit has already built.

Architecture layer Loop stage OBSERVE 1. Physical layer IoT devices, CCTV, mobile/tablet apps, RFID/QR/facial check-in OBSERVE 2. Ingestion API layer feeding a broker - MQTT for device telemetry, Kafka for event streams UNDERSTAND 3. Processing Stream processing enriches raw events into operational signals DECIDE 4. Decisioning - Rules Engine Deterministic rules first; ML-based prioritization layered on later ACT 5. Action - Workflow Engine Assigns an owner, notifies, and escalates on a timer if unresolved VERIFY 6. Storage & Analytics Existing warehouse (Redshift/Postgres) extended with real-time event tables VERIFY 7. Visibility - Command Centre Metabase dashboards, extended with a real-time tiered view: HQ → Regional → Cluster → Center continuous feedback loop, not a one-time build Already exists Tablets, CCTV, apps (2022 pilot) New for COIP Broker, stream processing, rules engine, workflow engine - none of this is documented as existing today Already exists Warehouse + Metabase (2020 write-up), extended here

12.1 High-Level Flow

  • Physical layer - IoT devices, CCTV, mobile/tablet apps, RFID/QR/facial check-in
  • Ingestion - API layer feeding a message broker (MQTT for device telemetry, Kafka for event streaming)
  • Processing - stream processing for event enrichment, feeding a rules engine
  • Decisioning - rules engine (initially deterministic; ML-based prioritization later)
  • Action - workflow engine handling assignment, notification, and escalation
  • Storage & Analytics - existing data warehouse (Redshift/Postgres), extended with real-time event tables
  • Visibility - Metabase-based dashboards extended with a real-time Command Centre layer

12.2 Precedent Technology Choices

Rather than proposing a generic stack, I want the architecture to explicitly build on what Cult.fit's engineering team has documented investing in and validating in production: Redshift/Postgres for the warehouse, Metabase for dashboards and alerting, S3 (with Redshift Spectrum) as a data lake for cold storage, Hevo as an ETL layer for third-party sources, and lightweight on-device ML (TFLite-class models) for edge inference where latency and cost matter - as already piloted in the facial-attendance system.

12.3 New Components Required

  • A message broker layer (Kafka/MQTT) for real-time device and CV events - not clearly present in the documented analytics stack, which is oriented around periodic and near-real-time (15-minute sync) batch data
  • A rules/workflow engine capable of stateful, multi-step escalation logic
  • CCTV-based vision-tech pipelines, extending the edge-inference pattern used for facial attendance to new camera-based use cases

13AI Opportunities

Comparable operators at much larger scale validate the direction of this roadmap. Domino's built its store-operations visibility on a real-time streaming platform (piloted with Talend and Apache Kafka/Confluent across roughly 25 U.S. stores from 2022) so that franchise owners and store teams could see order volume and efficiency metrics as they happened, rather than after the fact - the same architectural bet this case study proposes for center-level events. Kripakar Krishnamurthy, Domino's Director of Enterprise Data Engineering, described the goal in terms worth borrowing directly: giving store managers "a smart watch for their in-store operations" - reliable, live information about what's happening right now, not a historical report.

Domino's has since layered AI on top of that real-time backbone, including a generative AI assistant (built with Microsoft's Azure OpenAI Service) designed to help store managers with inventory, ordering, and staff scheduling, plus ML-based delivery-route optimization and order-readiness prediction models reported to have improved from roughly 75% to 95% accuracy - evidence that once a real-time data layer exists, the AI opportunities on top of it compound quickly.

OpportunityDescription
Predictive maintenanceUse equipment usage/failure history to flag machines likely to fail before they do
Anomaly detection in CV feedsFlag unusual occupancy, safety, or presence patterns without hand-coded rules for every case
Natural-language summarizationConvert dashboard and incident data into a plain-English daily brief for HQ leadership
Smart escalation prioritizationRank open incidents by likely member-experience impact, not just age
Root-cause clusteringGroup recurring incidents across centers to surface systemic (vs. one-off) issues

Given the fairness issue Cult.fit's own team already identified in its facial-recognition work, I think any new AI/CV capability under COIP needs an explicit accuracy-and-bias evaluation step before rollout, using a demographically representative internal dataset. Industry guidance on IoT/AI rollouts at scale is consistent on a second point too: anomaly-detection models should be introduced deliberately and tuned for low false-positive rates, since alert fatigue - not lack of data - is the most common reason frontline teams stop trusting an operational AI system.

14MVP

Industry pattern across retail, QSR, and manufacturing IoT rollouts is consistent: start with one site (or one line, one store, one cluster), one clearly-painful problem, and a short pilot window - typically 6–8 weeks for a narrow pilot, up to 3–6 months for a fuller single-site pilot - before committing to a wider rollout. Domino's own real-time operations platform began as a single-purpose store-analytics use case before expanding into marketing and international scale. My MVP follows the same discipline: prove the operational-latency thesis on a single cluster and a single workflow before any large infrastructure investment.

Proposed MVP Scope

  • Extend the existing facial-attendance tablets to also power a simple "trainer running late" automated workflow - no new hardware required
  • A lightweight Center Health dashboard (built in or alongside Metabase) combining attendance data, a manually-triggered audit-defect log, and incident status
  • One automated workflow end-to-end: trainer-absence detection → task creation → Center Manager notification → time-bound escalation to Cluster Manager
  • Manual (non-CV) incident logging for equipment and cleaning issues, to validate the workflow-engine and escalation logic before investing in CCTV-based detection

Success Criteria for MVP

  • Reduction in median time-to-first-response for trainer-absence incidents in the pilot cluster vs. a matched control cluster
  • Cluster Manager and Center Manager qualitative feedback on reduced WhatsApp-based escalation
  • Data pipeline proven reliable enough to justify investment in CCTV-based detection for Phase 2

15Roadmap

I deliberately benchmarked the phase durations below against comparable real-world rollouts rather than assuming them: retail and manufacturing IoT programs typically run small pilots in 3–6 months and reach full multi-site scale only after 12–24 months, expanding in reusable "waves" that reuse the architecture proven in the pilot rather than redesigning it each time.

PhaseScopeDuration (indicative)
Phase 0 - PilotSingle cluster; trainer-absence workflow; Center Health dashboard MVP0–3 months
Phase 1 - Regional RolloutExpand workflow automation to 2–3 additional workflows (maintenance, cleaning); roll out to a full region3–7 months
Phase 2 - Vision-Tech ExpansionCCTV-based detection for occupancy and trainer presence in flagship centers7–12 months
Phase 3 - Full-Scale IoT + Command CentreSensor rollout across equipment; HQ-level real-time Command Centre live across all centers12–18 months
Phase 4 - Predictive LayerPredictive maintenance, anomaly detection, smart escalation prioritization18+ months

I'd have each wave deliberately reuse the prior phase's architecture (same event schema, same workflow engine, same dashboard layer) rather than re-platforming - this is the single biggest determinant of whether retail and manufacturing IoT rollouts stay on budget as they scale past the pilot site, per implementation guidance from comparable large-scale programs.

16Risks

RiskMitigation
CV model bias across skin tones / lighting (already observed internally)Mandatory fairness evaluation on representative datasets before any new CV launch
Member and staff privacy concerns around CCTV-based monitoringClear consent and data-retention policy; use aggregate/derived signals, not stored raw video, wherever possible
Frontline change management - new tools add friction for trainers/housekeepingDesign for near-zero extra taps; pilot with heavy frontline feedback loops before wider rollout
Cost of sensor/camera hardware at 700+ center scalePhase rollout by center tier/flagship status; prioritize software/workflow wins before hardware capex
Over-reliance on rules engine leading to alert fatigueStart with a small number of high-confidence workflows; expand only as false-positive rates are proven low - flagged as the top reason frontline teams abandon operational AI/IoT systems in comparable retail and manufacturing rollouts

17Success Metrics

At the platform level, I'd judge COIP's success against a small number of outcome metrics that tie directly back to the operational-latency thesis:

  • Reduction in median and P90 operational-issue time-to-resolution, cluster by cluster
  • Increase in Center Health Score and reduction in audit defect rate over time
  • Increase in the share of "verified" (sensor/CV-based) vs. self-reported operational data
  • Improvement in member NPS and reduction in complaint rate correlated with COIP rollout
  • Reduction in manual escalation volume (WhatsApp/ad-hoc) as automated workflows take over

AAppendix - Mapping AECC Experience to This Role

The domains differ, but the underlying product thinking transfers directly: both are about managing distributed, semi-autonomous human operations toward a consistent outcome, using data and workflow design rather than headcount.

AECC (Education Counselling)Cult.fit (Fitness Operations)
CounsellorTrainer / Center Manager
Meeting → Application → Offer → EnrollmentCheck-in → Class → Attendance → Renewal
Distributed counsellor performance visibilityDistributed center performance visibility
Conversion and retention optimizationRenewal and retention optimization
Manual follow-up workflowsManual escalation via WhatsApp/audits

This mapping - distributed operations, resource utilization, and conversion/retention optimization through structured product thinking rather than domain-specific fitness expertise - is the throughline I'd bring to this role: the domain is new to me, but the operating pattern isn't.