
Jonathan Fulton
Staff Software Engineer

Amy Zhou
Product Manager

Uday Tennety
Head of Technology Partnerships
ChatGPT Work has become a common starting point for data and product teams. Analysts open it to compare launch adoption across segments, diagnose a metric that moved overnight, or turn a week of scattered numbers into a readout that a leader can act on. But the moment teams ask whether their experiment actually caused an effect they’ve observed, the conversation stalls. Someone has to leave the thread, open a separate tool, run the analysis, and paste the answer back into the channel, often losing the surrounding context along the way.
We’re pleased to announce the Datadog Experiments plugin for ChatGPT Work, which brings your experiment results into ChatGPT alongside the rest of your business data. Once the plugin is enabled, anyone on your team can ask about a running or completed experiment in plain language. The answer they get is grounded in Datadog’s statistical engine, with the same building blocks your experimentation program already relies on, including guardrails, methodologies, and links to connected context in Datadog.
In this post, we’ll walk through using the Datadog Experiments plugin to:
Combine experiment results with warehouse and business context
Catch experiment health problems before they cost you a rerun
Ask about experiment results without leaving your workflow
Historically, the bottleneck for experiment readouts has been whoever is fluent in the analysis tool. A product manager wants to know whether the new onboarding flow moved activation; a support lead wants to know whether a pricing test changed ticket volume. Both file a request, and both wait.
The Datadog Experiments plugin removes that queue. Referencing @Datadog Experiments in ChatGPT lets you query the same results surfaced in Datadog: lift and confidence intervals per variant, sample sizes, the metrics under test, and the guardrail metrics running alongside them. Because Datadog Experiments computes results continuously as an experiment runs, the answers reflect current data rather than a nightly snapshot.
“@Datadog Experiments, how is the checkout redesign test tracking against conversion and p95 latency? Break the results out by device type and new versus returning visitors.”
This is particularly helpful for questions regarding segmentation. A flat top-line result often hides a strong win in one cohort and a regression in another, and exploring these differences in a conversation is faster than building a dashboard for each hypothesis. The Datadog Experiments plugin returns segment-level results with the built-in variance reduction and sequential testing already applied, so peeking at intermediate results doesn’t inflate your false-positive rate the way ad hoc mid-flight checks do.
Every response cites the underlying experiment and links back to Datadog, so a promising finding in ChatGPT is one click away from the full analysis and from Session Replay and Real User Monitoring (RUM) data for the sessions behind it.
Here is an example experiment result from the plugin, generated by prompting ChatGPT to chart the results of an experiment:

Combine experiment results with warehouse and business context
Readouts rarely get questioned because of the statistics. More often, it’s because the metric in the experimentation tool doesn’t match the metric that the finance team reports. Datadog Experiments addresses this by measuring impact against source-of-truth business metrics in your native data warehouse and leverages the same definitions your business already runs on.
Inside ChatGPT, that means experiment results sit next to the rest of your data stack. You can pull experiment lift from Datadog Experiments, revenue from your warehouse, and deal context from your CRM in a single prompt, then have ChatGPT assemble the comparison:
“Compare day-30 retention lift from our last three onboarding experiments in @Datadog Experiments against ARR by segment from @Databricks Genie, and draft a summary of which cohort we should build for next.”
OpenAI’s data solution connects business data with company context, helping you identify what changed and why, and then turns that insight into a decision. The Datadog Experiments plugin adds the crucial piece: causality. Without the plugin, ChatGPT can tell you that a metric moved and correlate that move with changes that happened at the same time. With the plugin, ChatGPT can tell you which change caused the move, with a confidence interval attached, and deliver it wherever the next decision gets made: for example, an executive readout, a product requirements document (PRD) for the follow-up work, or a Slack summary for the team that shipped the change.
Catch experiment health problems before they cost you a rerun
A restarted experiment costs you in ways that don’t show up in a bill. You burn weeks of traffic, the roadmap slips, and the team’s confidence in the program takes a hit, usually because of a sample ratio mismatch or an instrumentation bug that nobody caught until the readout looked strange.
Datadog Experiments continuously validates experiment health, automatically detecting sample ratio mismatch, traffic imbalance, and instrumentation issues while a test is running. The plugin surfaces those signals in ChatGPT, too, so a health problem shows up in the same place you asked about the results.
“@Datadog Experiments, flag any active experiments with sample ratio mismatch, traffic imbalance, or a guardrail regression this week.”
Because Datadog Experiments monitors latency, errors, and crashes alongside behavioral metrics, the same query surfaces performance regressions introduced by a variant.
Keep AI-assisted analysis governed and auditable
Opening experiment results to everyone who can type a question raises fair concerns: Who can see what, and can you trust the answer?
The plugin inherits access controls on both sides. ChatGPT Work governs data access with role-based access control (RBAC), and the plugin respects your existing Datadog permissions so a teammate only retrieves experiment data they could already view in Datadog. Datadog records every query the plugin makes in Audit Trail, giving you the same visibility into AI-driven access that you have across the rest of your organization’s Datadog usage.
Just as importantly, the analysis isn’t a black box. Datadog Experiments gives you full visibility into methodologies, metrics, and underlying logic, and the plugin cites the experiment behind every claim.
Turn experiment results into decisions faster
Experimentation programs stall for organizational reasons more than technical ones: Results take too long to reach the people making decisions, and by the time they arrive, the team has already decided without them. The Datadog Experiments plugin for ChatGPT Work shortens that path by putting trusted, guardrailed experiment results where your team is already working, next to the business context that makes them actionable.
To learn more, see the Datadog Experiments documentation.
To start analyzing your Datadog experiments in ChatGPT, sign up for a 14-day free trial.
