Building trust in regulated healthcare
Nomad Health exists to remove the barriers between clinicians and the bedside. As healthcare demand grows across the United States, hospitals face pressure to fill critical roles quickly while ensuring every clinician is qualified. The company’s platform helps clinicians, hospitals, and staffing agencies connect, credential, and place talent more efficiently. Built on Google Cloud, Nomad supports more than 500,000 clinicians and over 10,000 active jobs at any given time, creating a complex marketplace where placing the wrong clinician can affect patient safety and where trust is essential.
AI plays an increasingly important role in that mission. Nomad uses machine learning and generative AI to power clinician-job matching, recruiter messaging, job-detail extraction, and profile creation. The company launched its first predictive model in 2021, a recommendation graph in 2022, and its first LLM-powered applications in 2023. “AI has been central to how we improve clinician matching and marketplace efficiency for years,” says Ben Long, Staff Software Engineer at Nomad Health. “Today, it plays a critical role in helping us operate a marketplace at this scale.”
As AI became more deeply embedded across the platform, the team encountered a new challenge. The systems were creating value, but they were also becoming harder to understand and more costly to operate. Engineers struggled to investigate hallucinations, identify the source of unexpected outputs, and manage workflow costs driven by retry behavior and non-deterministic responses. “We were moving from semi-deterministic systems into fully non-deterministic workflows,” says Long. “That changed how we had to build, test, and reason about these applications.”
For Nomad, visibility wasn’t simply an engineering concern. Every recommendation and automated workflow needed to be trustworthy, and the cost of running them manageable. The team needed to understand how its AI behaved, and what it cost, before scaling across the healthcare marketplace.
Ruling out a security incident
To better understand and manage its growing portfolio of AI-powered workflows, Nomad adopted Datadog Agent Observability. The platform gives engineers visibility into model behavior, response quality, workflow execution, and failure patterns across development and production. Instead of treating AI systems as black boxes, engineers can investigate why a model behaved a certain way, trace the factors that influenced its response, and improve future behavior. “As our AI systems became more sophisticated, the hardest part wasn’t building them. It was understanding why they behaved the way they did,” says Long.
One incident in particular demonstrated the value of this visibility. Nomad had built an AI-powered workflow that generated suggested messages recruiters could send to clinicians. During user acceptance testing, a recruiter flagged a generated message that contained suspicious, phishing-like language. Because the platform operates in a highly sensitive environment, the team immediately treated the issue as a potential security concern. “At first glance, it looked like something much more serious,” says Long. “We needed to determine whether this was a prompt injection attack, a security issue, or something else entirely.”
Nomad quickly assembled a response team, disabled the message-generation workflows, and began investigating whether it represented a broader vulnerability. Using Datadog Agent Observability, engineers traced the complete lineage of the response and quickly identified the source. Rather than finding evidence of prompt injection or a man-in-the-middle attack, they discovered that the problematic language originated directly from the foundation model itself. “Within minutes we could see where the response originated, which let us focus on fixing the problem instead of chasing a security incident that didn’t exist,” says Long.
The visibility allowed Nomad to determine the scope of the issue, rule out a broader vulnerability, and safely re-enable the workflow the same day. Just as importantly, the investigation gave business stakeholders confidence that the issue was understood and under control.
“Within minutes we could see where the response originated, which let us focus on fixing the problem instead of chasing a security incident that didn't exist.”
The incident also became an opportunity to strengthen the company’s AI governance. Nomad introduced additional evaluation frameworks and Datadog runtime evaluations with reporting, improving quality assurance and creating auditable safeguards around AI-generated content for a regulated environment.
Maintaining cost-efficiency across every AI workflow
As Nomad’s use of AI grew, managing the cost of these workflows became as important as ensuring their quality. Retry behavior and non-deterministic responses were driving up the cost of individual interactions, limiting how widely the team could deploy them. With Datadog Agent Observability, engineers could understand the operational costs of each AI workflow, identify limitations in external APIs, and improve their retry and dead-lettering strategies. It also extended into local and development environments, catching issues before production.
This cost visibility became even more important as Nomad evolved from a services-focused organization into a software platform serving hospitals and staffing agencies. The transition introduced new complexity around cost management, and understanding what each AI workflow costs to operate became essential to delivering enterprise-ready products at scale. “As we transition from a services business to a software platform, understanding the cost of every AI workflow becomes a business requirement, not just an engineering concern,” says Long.
Without that visibility into cost, Nomad could not have made the optimizations needed to bring the platform to market.
“As we transition from a services business to a software platform, understanding the cost of every AI workflow becomes a business requirement, not just an engineering concern.”
Scaling healthcare AI with confidence
The same observability practices now support Nomad’s broader platform. Beyond Agent Observability, the team relies on Datadog tracing, logging, monitors, alerts, error tracking, Kubernetes observability, Real User Monitoring, and SLO monitoring to maintain reliable experiences across its Google Cloud environment. That unified visibility has also helped beyond AI, surfacing latency spikes and slow queries affecting job search.
That same trace-level visibility, alongside the cost visibility, gave Nomad confidence its platform was enterprise-ready before serving other healthcare organizations.
Today, Nomad identifies issues earlier, deploys with greater confidence, and accelerates development across its AI-powered products. The combination of Google Cloud and Datadog helped Nomad maintain 99.9%+ uptime while supporting more than 1,000 production deployments in 2025. With greater visibility into how its AI systems behave and what they cost, teams spend less time guessing and more time improving the experiences that help clinicians discover opportunities and healthcare organizations fill critical roles. “Every recommendation, message, and workflow we automate has to earn trust,” says Long. “Datadog Agent Observability helps us understand how our AI systems behave so we can scale responsibly across healthcare staffing.”
“Every recommendation, message, and workflow we automate has to earn trust. Datadog Agent Observability helps us understand how our AI systems behave so we can scale responsibly across healthcare staffing.”