VROONG unifies observability and builds a FinOps culture to support real-time logistics at scale | Datadog
VROONG unifies observability and builds a FinOps culture to support real-time logistics at scale

case study

VROONG unifies observability and builds a FinOps culture to support real-time logistics at scale

About VROONG

VROONG is a delivery platform that leverages advanced technologies such as artificial intelligence and big data to provide premium, customized delivery services tailored to businesses and customers.

Delivery & Logistics
200+ Employees
South Korea
“Datadog is more than just a monitoring tool—it's a powerful solution that empowers everyone to take ownership of FinOps.”
“Datadog is more than just a monitoring tool—it's a powerful solution that empowers everyone to take ownership of FinOps.”
Lee Seung-Yoon SRE Chapter Team Member VROONG

Why Datadog?

  • Consolidates multiple tools into a single platform
  • Provides end-to-end visibility across distributed services
  • Accelerates root cause analysis with correlated telemetry
  • Visualizes Kubernetes architecture and service dependencies
  • Enables real-time pod resource monitoring without added setup
  • Surfaces cloud cost data directly in engineering workflows
  • Attributes costs by service and team to drive accountability
  • Detects cost anomalies before spend escalates

Challenge

VROONG’s fragmented monitoring environment required engineers to switch between tools, slowing incident response and limiting visibility across services.

Key Results

50% reduction

In average daily cloud costs

Unified monitoring

Across metrics, traces, and logs

Established a FinOps culture

With shared cost ownership

Faster root cause analysis

With end-to-end visibility

Fragmented monitoring increases operational complexity

VROONG is a logistics platform that enables restaurants, retailers, and brands across South Korea to offer fast, reliable delivery—including same-day and on-demand fulfillment. Using AI and big data, the platform optimizes real-time routing, driver allocation, and delivery efficiency, often getting orders to customers within the hour. Behind that speed is a large-scale, distributed system managing orders, drivers, and inventory simultaneously across the country.

As the platform scaled, VROONG’s engineering team managed infrastructure across multiple disconnected monitoring tools. Investigating an incident required switching between systems to piece together a root cause, slowing collaboration and adding overhead. Alert formats and policies were inconsistent across teams, and siloed application and infrastructure metrics made it difficult to understand how services interacted. The team needed a unified view across its infrastructure—not just to respond to incidents faster, but to operate more efficiently at scale.

VROONG delivery service

Cutting through complexity with unified observability

VROONG consolidated its observability data onto Datadog, bringing metrics, traces, and logs into a single platform. With unified dashboards spanning Application Performance Monitoring (APM), infrastructure, and logs, engineers can move from a high-level view of system health to detailed diagnostics without changing tools—and see how issues propagate across services.

The value of that end-to-end visibility became clear during a recent database incident. After detecting a CPU spike, an engineer started from a high-level dashboard and drilled into Database Monitoring, where query samples surfaced a long-running REINDEX operation as the root cause. The team could also see how the same query was driving latency increases in critical SELECT queries downstream. With both the cause and its broader impact visible in one place, they were able to take targeted action and restore performance significantly faster than before.

The team also adopted Datadog Notebooks to bring more structure to post-incident work, using them to document retrospectives and run weekly service review meetings. Centralizing links to relevant monitoring views in a shared workspace improved knowledge sharing and made post-incident analysis more consistent across teams.

Improving network visibility in Kubernetes

One of the team’s more significant wins was eliminating a major visibility gap in its network layer using Datadog Cloud Network Monitoring (CNM). By collecting real-time data on how services communicate across their containerized environment, VROONG’s engineers can now visualize network architecture and service dependencies at a glance. In Kubernetes, this helps teams assess incident impact more quickly and respond with greater confidence.

“Datadog is more than just a monitoring tool—it's a powerful solution that empowers everyone to take ownership of FinOps.”

CNM’s service map has also become a practical onboarding tool, helping new engineers understand system architecture without relying entirely on tribal knowledge. Kubernetes resource monitoring is now more accessible as well—with real-time visibility into pod resource usage available without additional setup, teams can evaluate configurations, prevent performance degradation, and plan capacity ahead of traffic spikes.

Driving a FinOps culture across engineering and operations

With a stronger observability foundation in place, VROONG turned its attention to cost. In Kubernetes environments, shared resources and idle capacity make it difficult to attribute costs accurately to specific services—a challenge that grows with scale. The team set out to build a FinOps practice where cost awareness is shared across engineering, operations, and planning.

“True FinOps isn't achieved when it's ‘someone's responsibility.’ It means development, operations, and planning all being aware of costs and collectively driving improvement.”

Making cost data visible within the same dashboards engineers already used for monitoring helped drive that shift. Teams began optimizing resources earlier in the development lifecycle once they could see how changes affected cloud spend. In one case, average daily pod costs dropped from $14 to $7, contributing to broader cost reductions across the platform. A feedback loop emerged: cost insights surfaced, developers acted, and results were validated in the same platform. “Datadog is more than just a monitoring tool—it’s a powerful solution that empowers everyone to take ownership of FinOps,” says Lee Seung-Yoon, SRE Chapter Team Member at VROONG.

Tagging by service and team improved cost attribution and surfaced optimization opportunities that had previously been difficult to identify. Network visibility also revealed hidden costs, such as cross-zone data transfer and per-workload network usage, making them accessible to engineers without requiring specialized expertise. “True FinOps isn’t achieved when it’s ‘someone’s responsibility,’” says Seung-Yoon. “It means development, operations, and planning all being aware of costs and collectively driving improvement.”

Automating cost optimization

VROONG’s next step is to automate more of its cloud cost management workflows. In the near term, the team plans to use cost anomaly detection monitors with attached dashboards, enabling engineers to investigate cost spikes without switching tools.

Further out, the company is exploring automated workflows that analyze metrics, generate summaries of cost anomalies, and surface findings directly in collaboration tools like Slack or Notion—shifting from reactive investigation to more proactive optimization.

The long-term goal is to calculate a “cloud cost per delivery” metric by combining APM, infrastructure, and cost data with business metrics. Linking infrastructure spend directly to each delivery would provide clearer insight into core operational costs and support more data-driven decision-making across the organization.

Resources

gated-asset/gartner2025obplatformopengraph200x630

guide

2025 Gartner® Magic Quadrant™ for Observability Platforms
gated-asset/reduce_it_costs_ebook_og

ebook

Reducing IT Costs with Observability