17 min read

Source: Roblox Creator Hub · CC BY 4.0 · View source · Code samples: MIT Imported 2026-10-03. Formatting adapted for this site.

Experiments

Experiments let you run in-game and matchmaking A/B tests to measure the causal impact of changes to your game. For example, you can show different onboarding experiences to different players and measure the difference in playtime, retention, and other key performance indicators.

Experiments are excellent for measuring the following:

Overview of the Experiments page on Creator Hub

Create experiments

Experiments come in two types:

In-game

  1. If you don't already have a config, create one for your game.

  2. On the Creator Hub Experiments page for your game, click Create experiment.

  3. For Type, choose In-experience.

  4. Specify a name, goal metric, and planned duration for the experiment. Experiments run for between 14-60 days.

    Regardless of what you choose as your goal metric, experiments track all metrics in the list.

  5. Choose a percent rollout. This number is the percentage of players that you want to include in the experiment.

    In general, the more people you include in an experiment, the better the data, but use your judgment on what's best for your game.

Note

Pay attention to the minimum detectable effect, which is the smallest change in the selected metric that the experiment can reliably detect.

  1. Specify variants and percentages.

    Variants are alternative values for your config. For a numeric config key bossHealth with a control value of 500, you might specify a variant of 300. You can have up to two variants and one control in an experiment.

    Percentages dictate how to assign variants within the experiment rollout. Consider the following example:

    • You choose an overall rollout of 40%.
    • You specify two variants and a 50/50 split between them and the control.

    In this example, 60% of your users are excluded from the experiment; these users receive the control and have no impact on experiment results. Approximately 20% of your users receive the control as part of the experiment. Another 20% receive the variant. Depending on your player count, this distribution might not be large enough to yield actionable results.

Note

When in doubt, we recommend simple 50/50 experiments of one variant and one control on large rollouts. The experiments are easier to configure, and the results are easier to interpret.

Variant page

  1. (Optional) Target the experiment to a specific audience so that only matching players are eligible for enrollment. Targeting uses the same player attributes as conditional configs, and you can copy existing conditions directly from your configs. For more information, see Target experiments to specific audiences.

    The optional targeting step during experiment creation

  2. The final step is scheduling. You can start experiments immediately or schedule them for a later date and time. After you schedule an experiment, you can't change its configuration (duration, rollout percentage, variants, etc.), but you can reschedule it.

Matchmaking

  1. If you don't already have a custom matchmaking configuration, create one for your game.

  2. On the Creator Hub Experiments page for your game, click Create experiment.

  3. For Type, choose Matchmaking.

  4. Specify a name, goal metric, and planned duration for the experiment. Experiments run for between 14-60 days.

    Regardless of what you choose as your goal metric, experiments track all metrics in the list.

  5. Choose a percent rollout. This number is the percentage of players that you want to include in the experiment.

    For matchmaking experiments, we highly recommend 100% rollout. It minimizes the risk of isolating players and leads to faster, more reliable results.

Note

Pay attention to the minimum detectable effect, which is the smallest change in the selected metric that the experiment can reliably detect.

  1. Choose the places you want to include in the experiment, and specify variants.

    Variants are alternative matchmaking configurations. You can include up to three variants in matchmaking experiments. Regardless of how many variants you include, Roblox splits players equally between all variants.

Note

When in doubt, we recommend simple 50/50 experiments of one variant and one control on large rollouts. The experiments are easier to configure, and the results are easier to interpret.

Matchmaking experiment variant page

  1. The final step is scheduling. You can start experiments immediately or schedule them for a later date and time. After you schedule an experiment, you can't change its configuration (duration, rollout percentage, variants, etc.), but you can reschedule it.

Metrics

Experiments track all of the following metrics over the experiment duration. These metrics are updated every 24 hours while the experiment is running. However, during the first 24 hours of your experiment, Playtime, Payer conversion rate, and ARPU also update every 5 minutes so you can catch critical regressions (early harm) quickly.

MetricDescriptionEarly harm?
D1 retentionPercentage of new players who returned to your game after one day.No
D7 retentionPercentage of new players who returned to your game after one week.No
PlaytimeAverage amount of time players spent within your game. Cumulative for the duration of the experiment.Yes
ARPUAverage revenue per user. Revenue divided by the number of players. Cumulative for the duration of the experiment.Yes
ARPPUAverage revenue per paying user. Revenue divided by the number of players who purchased a game-related item. Cumulative for the duration of the experiment.No
Payer conversion ratePercentage of players who purchased a game-related item. Cumulative for the duration of the experiment.Yes
Session timePlaytime divided by number of sessions. Cumulative for the duration of the experiment.No

Experiment status

The Experiments page shows the following statuses for experiments.

StatusDescription
CompletedThe experiment is over, which happens when you stop it manually, when you reach a decision, or automatically shortly after the decision date (14 days after for in-game, immediately for matchmaking). You can still review the details and results.
Decision neededThe experiment has reached its decision date. Now is a good time to review the results.
RunningThe experiment is running but has yet to reach its decision date.
ScheduledThe experiment is scheduled to start at a future date.
DraftThe experiment hasn't been started or scheduled. You can finish setting it up.

Target experiments to specific audiences

By default, an in-game experiment can enroll any player in your rollout percentage. Targeting lets you run the experiment on a specific audience instead, such as players in certain countries, specific tenure windows, or active spender tiers. Targeting uses the same attributes as config targeting.

Note

Targeting only applies to in-game experiments.

Set up targeted experiments

You configure targeting criteria when you create the experiment. After the experiment starts, you can't change the targeting rules, the control value, the variants, or any conditional rules that the config key uses.

You can reuse existing conditional rules from your configs to define your targeting, but the experiment stores an independent copy of each rule. Editing or deleting the original config condition later has no effect on the experiment's targeting.

Because narrowing your audience reduces your sample size, Roblox updates the minimum detectable effect (MDE) using an estimate of the audience that matches your targeting criteria. Make sure your targeted audience is large enough to produce statistically meaningful results.

How Roblox resolves targeted values

Roblox evaluates experiments first, before any rules from standard conditional configs. When a config key has an active experiment, Roblox resolves the value in this order:

  1. The active experiment value, if the player is targeted and enrolled.
  2. The first matching config condition.
  3. Any subsequent matching config conditions.
  4. The default value.

Players in the control group receive the value that your existing config rules produce. You can't run more than one active experiment on the same config key at a time.

Per-session evaluation

Roblox evaluates targeting attributes per session, so an individual player's eligibility can change over time:

A player's data stays attributed to the variant (or control) they were enrolled in, even if they later stop matching the criteria. For example, D7 retention still counts a player who enrolled on day 0 and whether they returned on day 7, even if they no longer qualify by day 4.

Config changes while an experiment runs

To prevent configuration drift and corrupted experiment data, Roblox locks the config key the moment an experiment starts running:

The lock releases when the experiment completes.

Add experiments to your code

Note

This section only applies to in-game experiments. Matchmaking experiments do not require code changes.

Applying in-game experiments is similar to applying configs. The main difference is the use of ConfigService:GetConfigForPlayerAsync() rather than ConfigService:GetConfigAsync().

GetConfigForPlayerAsync() retrieves a player-specific snapshot. When you call GetValue(), the snapshot checks for an active experiment and enrolls (or doesn't enroll) the user based on the rollout percentage.

local ConfigService = game:GetService("ConfigService")
local Players = game:GetService("Players")

local function onPlayerAdded(player)
    local playerConfig = ConfigService:GetConfigForPlayerAsync(player)
    local leaderboardColor = playerConfig:GetValue("leaderboardColor")
end

Players.PlayerAdded:Connect(onPlayerAdded)

Note

Wait to call GetValue() until you need it. Calling GetValue() too early can cause you to enroll players who never interact with the part of the game you're experimenting on.

Custom enrollment

Note

To target players by attributes like country, tenure, or payer status, use experiment targeting instead. The approach in this section is for in-game state, such as how far a player has progressed.

If you want to enroll only players that meet criteria based on in-game state, you have to write additional code to check for those criteria and only then call GetValue() to enroll them in the experiment. Consider the following example:

Your code might look something like this:

local function getControlScheme(player, racesWon)
    if racesWon < 20 then
        return "standardScheme"
    else
        -- Player has many wins, enroll in experiment
        local playerConfigSnapshot = ConfigService:GetConfigForPlayerAsync(player)
        if playerConfigSnapshot:GetValue("useNewControlScheme") then
            return "newScheme"
        else
            return "standardScheme"
        end
    end
end

If you want the control scheme to persist on subsequent sessions, you likely need to add a value to the player's entry in a data store.

View and interpret results

Click View to see details and results. In the Details And Progress tab you can see the total number of players enrolled, as well as the number of players that received the control value and each variant. Viewing this page early in the experiment is useful strictly for making sure the experiment is running properly, not for taking action. Before taking action, see Best practices.

The details page for an experiment

Early harm metrics

For the first 24 hours, or until the first daily results land, early harm metrics update every 5 minutes so you can catch unintended critical harm early. Click Metrics to see these results. Look for critically harming metrics highlighted in red.

Hover over a metric and click View confidence to see the confidence interval.

Early Harm Results for an experiment

During this window, results test only for harm, so the confidence interval has no lower bound, only an upper bound. A metric is flagged as critically harmful only if the entire confidence interval falls below the critical harm threshold. In the following example, Playtime is at -12.67% with an upper bound of -2.73%. Even though the estimate is lower than the critical harm threshold of -10% and the entire confidence interval is below 0%, the metric is not classified as critically harmful because the interval is not entirely below the harm threshold.

Confidence interval for an Early Harm metric

Use these results only to guide early stopping decisions in the first 24 hours. Do not use them for launch decisions, rely on daily results instead.

Early harm thresholds

Each metric has its own early harm threshold. If a variant's lift falls below this threshold, it will be classified as critically harmful to that metric. However, for metric variant comparisons with fewer than 10,000 players enrolled across the variant and the control, variant lifts are shown but decisions on harm will not be made until the sample size provides enough data to reliably detect harm.

MetricEarly harm threshold
Playtime-10%
ARPU-20%
Payer conversion rate-20%

Early harm notifications

If critical harm is detected in your experiment during the early harm analysis period, a notification is delivered to the game owner and experiment owner for early action. Notifications are delivered via email, Creator Hub notification tray, and through an optional webhook. Upon receiving an early harm notification, you should review the experiment metrics and decide whether or not the experiment should be stopped early.

Early Harm creator hub notification

Daily results

After an experiment has run for at least 24 hours, click Results to see the latest results, which update every 24 hours. Look for statistically significant changes in goal metrics, which the dashboard highlights in green or red. These changes are more likely to show the impact of your variant and less likely to be false positives or negatives.

The details page for an experiment

A metric is statistically significant when the confidence interval for its percent change does not overlap with 0%. In the following example, ARPU is up 2.37%, with lower and upper bounds 0.55% and 4.19%, which makes the change statistically significant.

Confidence interval for a metric

For convenience, the results page lets you replace the default config value with one of the variants from the experiment.

Sample ratio mismatch

If your experiment is failing to enroll users into the variants in the expected proportions, an alert banner for Sample Ratio Mismatch (SRM) will appear in the experiment results tab. If SRM is detected, it is recommended to stop and restart your experiment.

SRM is checked every 5 minutes for the first 24 hours of your experiment, and then daily after. SRM can invalidate experiment results; do not make launch decisions based on compromised and unreliable results.

Sample Ratio Mismatch banner

Make a decision

When your experiment concludes, click Make decision to start a guided rollout. Roblox shows any warnings about statistical significance and experiment duration, then prompts you to select a winning variant or keep the control. Based on your choice, Roblox rolls out the appropriate change to your permanent configs; returning to the Configs page shows the new value. Click Change winner if you change your mind.

If you have unresolved staged changes on the config when you complete the experiment, Roblox temporarily stashes them and restores them on a best-effort basis after the rollout.

Your decision determines the config changes that Roblox proposes:

Best practices for experiments