Operating a global product with a small platform team
Delightroom is a Seoul-based wellness company on a mission to help users start their mornings successfully. Its flagship product, Alarmy, is recognized as the world’s #1 alarm app, with 82 million cumulative downloads and users in 97 countries, 85% of whom are outside Korea. Alongside Alarmy, the company operates DARO, a B2B ad monetization platform for mobile developers. Together these products generate more than $50 million in annual revenue.
At this scale, reliability is directly tied to business outcomes. Alarmy is used by millions of people around the world to start their day, so even small disruptions can affect user trust, engagement, and revenue. Supporting that experience requires systems that operate reliably with minimal manual intervention. “For us, reliability is not just a technical metric—it directly impacts whether our users can start their day as expected,” says Andy Kim, Foundation Group Lead at Delightroom.
Despite this scale, both products run on shared infrastructure managed by a small team of platform engineers. This operating model reflects a deliberate approach: invest in systems and tooling that scale with the business, allowing a small team to support global usage reliably. “We operate a global product with a very compact team, so everything we build has to scale operationally,” Kim explains.
“For us, reliability is not just a technical metric—it directly impacts whether our users can start their day as expected.”
To support that model, the Foundation Group, which is responsible for platform reliability, focuses on key service metrics such as availability, mean time to detection (MTTD), mean time to resolution (MTTR), and Apdex. At the same time, the team emphasizes operational efficiency, increasing the share of work that can be handled with minimal manual intervention. For a globally distributed, always-on product, observability plays a central role in achieving both goals.
Unifying observability across a shared platform
Delightroom needed a way to connect observability data across its shared infrastructure. Alarmy and DARO run on a shared Amazon EKS cluster with six core services, but their operational characteristics differ significantly. Alarmy traffic follows global wake-up patterns, with demand shifting across time zones, while DARO operates in a latency-sensitive advertising environment where bid response times are critical.
As the platform evolved, Delightroom focused on improving how observability data is connected across systems. Diagnosing issues required correlating signals from multiple sources. To simplify this process, Delightroom deployed Datadog across its entire stack, creating a unified observability layer that brings these signals together. “Before, the signals existed, but the challenge was connecting them into a single, actionable view,” says Kim.
Mobile RUM provides visibility into Alarmy’s iOS and Android applications, tracking crash-free sessions, ANR rates, view performance, and app startup time. RUM for web extends coverage to the Alarmy payment flow, the DARO dashboard, and internal tools.
“Before, the signals existed, but the challenge was connecting them into a single, actionable view.”
At the same time, visibility extends across the rest of the stack. Ingress monitoring provides traffic and error-rate insights at the edge, Application Performance Monitoring (APM) traces all core services end to end, and Kubernetes monitoring tracks container health, resource usage, and restart patterns. At the data layer, Amazon Aurora is monitored with anomaly detection that accounts for expected traffic patterns throughout the day.
Together, these signals feed into seven production service level objectives (SLOs), each targeting 99.9% to 99.95% over a 30-day rolling window. Synthetic Monitoring contributes uptime data directly to these SLOs, giving the team a continuous view of reliability across both products.
For Alarmy, this visibility maps directly to user experience. Users do not think in terms of system metrics—they care whether their alarm works as expected. RUM helps the team detect regressions in that experience early, before they affect a broader portion of users.
Alerts are structured to provide clear, actionable context, including impact, thresholds, and affected services, along with links to runbooks and dashboards. As a result, engineers can move directly from detection to investigation. “We design alerts to take engineers from signal to action in one step, without needing to stitch context together manually,” Kim adds.
Growing operations without increasing team size
With unified visibility in place, Delightroom focused on improving operational efficiency without expanding the platform team. The company supports two products generating more than $50 million in annual revenue with a platform team of just two to three engineers, and it has maintained this structure as the business has grown.
To further improve efficiency, Delightroom incorporated Bits AI into its incident response workflows. Bits AI supports investigation by generating likely root cause hypotheses based on real-time telemetry, allowing engineers to focus on higher-level decisions such as assessing impact and selecting mitigation strategies. As a result, the time required for detection and root cause analysis has decreased from approximately 20 minutes to about 5 minutes for many incidents. “With Bits AI SRE being on-call 24/7 for us, MTTR for our services has improved significantly,” says Kim. “Overall, we’ve seen MTTR improve by approximately 50%, while many investigations are already completed before our engineers sit down and open their laptops.”
Building on this approach, the team is expanding its use of Datadog APIs, CLI, and MCP server to create AI-driven workflows that operate directly on real-time observability data. These workflows support tasks such as alert triage and root cause analysis, helping the team manage routine operational work more efficiently. The team has also increased the percentage of issues resolved at the first level by approximately 50%, reducing the need for escalations and allowing engineers to focus on higher-value work. “Our goal is to connect signal to decision and action as seamlessly as possible, including for AI-driven workflows,” Kim explains.
“With Bits AI SRE being on-call 24/7 for us, MTTR for our services has improved significantly. Overall, we've seen MTTR improve by approximately 50%, while many investigations are already completed before our engineers sit down and open their laptops.”
This operational model also extends to compliance. Monitoring, alerting, incident response, and log retention data are captured as part of normal operations, making it easier to provide the evidence required for SOC 2 audits without additional preparation. “SOC 2 became a non-event for us because the evidence was already in the system,” says Kim.
Building a platform that scales with the business
As Delightroom continues to grow, the team is extending this approach to support long-term scalability and operational consistency. This includes expanding its use of SLOs to incorporate user experience metrics from RUM, aligning infrastructure reliability more closely with end-user outcomes.
“The value of Datadog has compounded over time. It's evolving from an observability platform into a foundation for AI-driven operations that lets us scale the business without scaling the team.”
Bits AI adoption is also increasing gradually, with workflows introduced where they provide clear operational value. This measured approach allows the team to improve efficiency while maintaining reliability. “We are taking a deliberate approach to Bits AI, expanding where it clearly improves speed and accuracy in operations,” Kim says.
This evolution reflects a broader shift in how Delightroom approaches operations. Observability has become more than a monitoring layer—it is now a foundation for both human and AI-driven workflows that support the entire business. “The value of Datadog has compounded over time,” says Kim. “It’s evolving from an observability platform into a foundation for AI-driven operations that lets us scale the business without scaling the team.”