Is DDoS Testing Safe to Run Against Production?
Yes. A DDoS test can be run safely against production when it is authorised, carefully scoped, actively monitored, and built so the customer can stop it at any moment. It is not risk-free: production testing deliberately places controlled pressure on live, customer-facing systems. The objective is to keep that risk bounded while testing the environment that would actually face a real attack.
Red Button actively recommends testing production where the organisation is ready for it. A staging environment can be a reasonable first engagement, but it rarely reproduces the security configuration, behavioural mitigation baselines, legitimate user traffic, and incident-response processes that decide the outcome of a real attack.
So the question is not simply production versus staging. Before a production test starts, the organisation has to settle the test window, provider notifications, shared dependencies, monitoring readiness, change-control requirements, and who needs to be available while traffic is running.
KEY TAKEAWAYS
- Production testing gives the most realistic assessment of your protection, because it tests the configuration, the traffic patterns, and the operational processes that a real attack would meet.
- Testing production carries real operational risk. The objective is to control that risk, not to pretend it does not exist.
- Notification is mandatory and usually sets the date. The customer must tell the mitigation provider, ISP, cloud provider, and data centre the planned window and the maximum expected throughput.
- Red Button runs the attack; your team can stop it at any moment, and Red Button complies immediately, without discussion.
- Staging is a legitimate starting point, not a substitute. Where an organisation is not ready for production on the first engagement, production comes in a later cycle.
Why Test DDoS Protection in Production?
The strongest argument is the simplest one: production is the environment you actually need to protect.
Security teams routinely mirror production controls into staging, but the two are rarely identical. Differences in rules, routing, capacity, protection settings, or application configuration all change how an attack is detected and mitigated. Across more than 2,000 controlled DDoS simulations, Red Button’s experience is that a clean staging result should not be treated as evidence that production will behave the same way.
Four things in particular are only visible in production.
It validates the actual security configuration
Production testing challenges the controls that are genuinely protecting the live service.
A staging environment may hold a carefully replicated copy of the production rules, and small differences still matter: a mitigation rule that is absent, applied to a different hostname, set to another threshold, or left in monitoring mode rather than enforcement. Testing the real environment removes the assumption that the copy represents the system an attacker will meet.
There is a second-order effect worth naming. A problematic rule in production is usually one nobody has noticed – if it were obvious, it would already have been fixed. Those are exactly the faults a test exists to surface.
It tests protection alongside real user traffic
Most staging environments carry little or no genuine user traffic.
That matters because protection controls do not operate in isolation during an attack. Rate limits and behavioural mitigation have to separate malicious traffic from legitimate users at the same moment, under the same pressure.
There are secondary effects too. When a service slows, real users retry failed requests, adding load on top of the attack while mitigation is already responding. That behaviour is very hard to reproduce in staging.
A production test can therefore expose false-positive risk – protection blocking your own customers – that an otherwise successful pre-production exercise leaves invisible.
It validates behavioural mitigation baselines
A growing share of DDoS protection is behavioural.
Cloud WAFs, scrubbing services and rate-limiting engines learn from historical traffic: normal request rates, user geographies, session shapes. Those baselines are built from the live environment. A staging system without that history does not necessarily trigger at the same point or make the same decisions.
If the question is whether a behavioural control will identify and respond correctly under real conditions, production is the only environment that carries the baseline being tested.
It tests the incident response process
A simulation should validate more than the mitigation technology.
In a production test, real monitoring generates the alerts, SOC and NOC teams respond, escalation paths are followed, and external providers react as they would in a genuine event – the AWS Shield Response Team, Akamai’s SOCC, and their equivalents behave differently when the event is real.
That lets you answer questions a configuration review cannot:
- Did the alert reach the right person?
- Did the team know what to check?
- Was the escalation path clear?
- Did the mitigation provider respond as expected?
- Did anyone hold the provider’s emergency contact?
A staging exercise can rehearse a runbook, but everyone involved knows they are responding to a staging system. CISA’s joint guidance with the FBI and MS-ISAC frames DDoS readiness as detection, response, and recovery rather than mitigation capacity alone – and those are the parts only a live event, or a production test, actually exercises.
What Are the Risks of Testing Production?
- Real customer impact.
The service under test is live. If a protection control behaves unexpectedly, legitimate users can see latency, failed transactions, or a temporary loss of availability. - Approval takes longer than the preparation.
In regulated industries, permission may have to pass through security, compliance, risk, and legal before anyone technical weighs in, and some organisations must notify external providers or a regulator on top of that. Each reviewer works to their own timeline. A few weeks is normal; heavier environments stretch into months, by which point the infrastructure has often changed enough that parts of the test plan need rewriting. - Fewer options to fix things immediately.
A problem found mid-test may need a ticket, a formal approval, or a scheduled change before it can be corrected. In staging, the same adjustment can often be made on the spot. That is a genuine advantage of staging, and one reason some teams start there. - Shared and third-party dependencies.
The blast radius can extend beyond the hostname being targeted where other services share infrastructure, network paths, upstream providers, or application dependencies. Those belong in the scope document as exclusions, or as declared and accepted exposure.
These are reasons to plan a production test carefully. They are not reasons to treat staging as equivalent evidence.
What Makes a Production DDoS Test Controlled?
A controlled test begins well before the first packet is generated.
Notify the relevant providers
The customer is responsible for notifying the mitigation provider, ISP, cloud provider, and data centre before the test. Each needs the estimated window and the maximum expected throughput. This is mandatory rather than a courtesy – an unannounced simulation produces a real incident response to a simulated incident – and provider lead times are usually the main factor deciding when the test can actually happen.
Cloud environments add their own rules. AWS allows approved DDoS Test Partners to test within its published policy without separate prior approval for each standard test, against resources in an account subscribed to Shield Advanced; tests outside those conditions go through an exception process with its own lead time. AWS currently publishes limits including 20 Gbps and 50,000 requests per second, with packet-rate restrictions that depend on the resource being tested, and traffic may not originate from an AWS resource. Azure requires an approved partner and a public IP in a subscription you own that is already protected by Azure DDoS Protection; Red Button’s Azure testing is subject to limits of 2.5 Gbps and 750,000 packets per second. Red Button appears on both providers’ lists.
Provider requirements belong in scoping, not in the week before the test.
Get the authorisation in writing
Provider policy is not the same as permission to test your own service. A scope document should name the targets, the exclusions, the vectors, the traffic ceiling, the window, and the people accountable on each side. That document is what makes the test defensible afterwards, which is why authorisation is also what settles the legal question.
Choose the test window around people as well as traffic
A production DDoS test does not usually require a maintenance window in the sense of planned downtime. It does require a defined test window.
Red Button commonly runs production tests during the customer’s off-peak hours – for most organisations, somewhere between late evening and early morning in their own time zone. Low traffic is only one input. The window also depends on whether the SOC, SRE and QA teams can be present, on the organisation’s change-control process, and on the lead time external providers need.
Sometimes a daytime test is the deliberate choice. If one objective is to evaluate how the SOC responds while fully staffed, running when the team is on shift produces better evidence than picking the quietest hour of the night.
One rule sits above the rest: Red Button does not run a planned simulation while the customer is already dealing with a live incident.
Make sure monitoring is ready before traffic starts
The customer needs visibility into its own environment throughout the test – the mitigation provider’s dashboard, network infrastructure, application health, and server performance – measured against a baseline captured beforehand. Without that baseline, degradation becomes a judgement call made under pressure.
In practice, the customer reports what it sees, often by reading the metrics out on the call or sharing the monitoring screen. Read-only access can be provided where policy allows and makes things easier still, but it is not required, and most customers do not provide it. Red Button runs independent external availability monitoring alongside.
The combination is the point: the two sides see different things. Red Button sees how the target behaves from outside; the customer sees what is happening behind the protection layer.
Match the traffic profile to the test objective
A controlled test does not mean every vector starts low and climbs the same way.
Gradual escalation is right when the objective is to find the point at which mitigation activates without overshooting it. The rate rises in stages – for example, 900, then 3,000, then 6,000, then 10,000 requests per second – with both sides watching how the environment responds, and the result of one stage deciding whether the next runs.
Starting at full rate answers a different question. A Hit & Run scenario opens aggressively on purpose, because what is being measured is how quickly the mitigation provider detects and responds to a sudden attack. Ramping up gently would change the very behaviour under test. Real attacks rarely announce themselves politely.
So the control is not “always start small”. It is that the attack profile, the maximum rate and the purpose of each vector were agreed in advance, in writing. That agreed ceiling should also sit below both your provider’s policy limit and your own known capacity – and it is worth resisting the instinct to book the largest number available, because volume is rarely what breaks a defence. Misconfiguration is, and misconfiguration shows up well below the ceiling.
Who Can Stop a Production DDoS Test?
The arrangement is straightforward: Red Button runs the attack, and the customer’s team can stop it at any moment.
Red Button does not require the customer to nominate a separate individual as the designated stop caller. Stop authority sits with the customer’s team, and if the customer asks for the test to stop, Red Button complies immediately and without discussion.
Testing can be halted instantly from the Red Button platform – a single control ends every running process at once. Once a stop is initiated, Red Button confirms that attack traffic has ceased. The customer then checks its own environment and validates that the service has recovered before anything else runs.
It is important to keep stakeholders on the call bridge throughout the duration of the test. Automated thresholds catch what was predicted. A support queue filling up, a downstream partner calling, a replica falling behind in a way no dashboard was configured to catch – those reach a human first, which is why the acceptable impact threshold is a starting point rather than the whole mechanism.
When Does Testing Staging First Make Sense?
Red Button recommends production, but that does not mean every organisation has to begin there.
Some customers need to build confidence in the process before putting a live environment under controlled attack traffic, and some cannot yet get internal approval for production. In both cases, starting with staging is a sensible first step. A staging test lets teams get familiar with the engagement, watch the process work, practise communication between teams, and find obvious configuration problems.
What staging cannot establish is that the live environment will behave the same way.
Four things a smaller or mirrored environment can still show: whether a mitigation rule exists and triggers at all; the sequence in which layered controls activate; whether alerting reaches a human; and configuration errors – the wrong allowlist, a rule applied to the wrong hostname, a control left in monitoring mode.
Four things it cannot: absolute capacity limits; any control that learns its baseline from real traffic; scrubbing or diversion behaviour where staging does not follow the same path; and autoscaling behaviour under cost pressure.
Percentages observed in an undersized environment do not extrapolate either – topology, routing, CDN behaviour, and the protection product itself all differ.
Red Button’s preferred path for a hesitant customer is therefore simple: start with staging on the first engagement, then move to production once confidence in the controlled process has been established. Staging is a stepping stone to production validation, not a permanent substitute for it.
Production Testing Is About More Than Attack Volume
A common misconception is that DDoS testing is mainly about discovering how much traffic an environment can absorb.
Capacity can be part of the exercise, but resilience depends on far more than a maximum Gbps figure. A production test shows whether a rule activates when expected, whether behavioural protection separates attackers from users, whether the mitigation provider reacts quickly enough, whether monitoring detects the event, and whether the internal response process works.
The threat landscape points the same way. Precision attacks below 1 Gbps cause disruption through how a request is processed rather than how much traffic arrives, and the same logic runs through the HTTP/2 family Red Button tests routinely – Rapid Reset, Continuation Flood, and HTTP/2 low-and-slow. Newer research continues in that direction: the HTTP/2 Bomb, disclosed in 2026 by the California research firm Calif, achieves disruption at negligible bandwidth. Exposure to vectors of this kind was never a function of volume.
This is also why production DDoS testing should not be confused with conventional load testing. DDoS testing, stress testing, and penetration testing answer different questions: a load test asks how the application behaves under increased demand, while a simulation asks whether the controls meant to detect, mitigate, and respond to hostile traffic do what you believe they do.
“We Already Have Cloudflare” – What a Test Actually Validates
A simulation does not evaluate whether Cloudflare, Akamai, or any cloud provider can absorb volume. It evaluates whether your deployment of that product behaves the way you think it does.
Across Red Button’s engagements, 68% of the protection faults identified were rated severe or critical – in organisations that had already deployed DDoS protection. Those are Red Button’s own findings rather than an industry benchmark.
Common ones include rate limits keyed per source IP that a distributed attack sidesteps, controls left in monitoring mode rather than enforcement, non-web protocols routed outside the protected path, and an origin still reachable directly, which lets traffic bypass the CDN it was meant to pass through.
Protection and validation are different disciplines, and buying the first does not deliver the second.
Planning a First Production Test
A first test is safest when the sequence removes surprises, not when it avoids production.
- Settle the questions that determine what the test can find. Which controls are deployed, what is in and out of scope, and what impact is acceptable – the nine questions to work through beforehand cover this.
- Confirm provider policy and secure written authorisation. Approved-partner status, target eligibility, and agreed ceilings, before a window is booked.
- Notify the traffic path. Mitigation provider, ISP, cloud provider, and data centre, each with the window and the maximum expected throughput. Start early; this usually sets the date.
- Agree on the window around availability. Off-peak in your own time zone, with SOC, SRE and QA able to attend and change control cleared.
- Consider a separate cloud account or subscription for a first run. A sandbox with the same protection profile reduces exposure, though it does not remove dependencies on shared services, provider limits or authorisation, and may not reproduce the real traffic path. Worth doing where a cloud DDoS simulation is what you are validating.
- Run the vectors one at a time, with a green light between each, and real-user metrics watched against baseline throughout.
- Debrief while it is fresh and convert findings into owned remediation items with a retest date.
Frequently Asked Questions
Is it safe to run a DDoS test against production?
Yes, provided the test is authorised, scoped, and actively controlled. Production testing carries residual risk because it affects a live environment, but that risk is bounded through provider coordination, monitoring on both sides, agreed traffic profiles, immediate stop capability, and continuous communication while traffic is running.
Does a production DDoS test require a maintenance window?
Usually not. Red Button agrees on a defined test window rather than planned downtime, and production tests are commonly scheduled during the customer’s off-peak hours.
The exact timing depends on traffic patterns, SOC, SRE, and QA availability, change-control requirements, and provider notification lead times. A typical session runs around three hours for six vectors, or around six hours for twelve, though the real duration depends on the scope and on what happens during the test.
Who can stop the test?
The customer’s team can require an immediate stop at any point. The Red Button team halts traffic instantly from our test platform, confirms that attack traffic has ceased, and waits for the customer to validate recovery before continuing.
Do you need access to our monitoring dashboards?
No. During the test, you watch your own dashboards – mitigation provider, network elements, application, and server health – and report what you see, while Red Button runs independent external availability monitoring alongside. Sharing your screen during the session makes the picture clearer for both sides, and read-only access is better still where your policies allow it, but neither is required.
Does a DDoS test require installing an agent or granting network access?
No. No traffic-generation agent or appliance is installed in your network, and no access into your systems is needed. Test traffic is generated externally and arrives at your public endpoints over the internet, exactly as attack traffic would. What you provide is scope – public IPs and hostnames, exclusions, written authorisation, and named contacts for the window. Monitoring agents you already run are unaffected, and the same applies to on-premise and hybrid infrastructure.
Is staging enough if we are not ready to test production?
It can be a useful starting point. Staging lets the organisation get familiar with the process and can expose configuration problems. It cannot reproduce production security settings, behavioural baselines, legitimate user traffic, routing, dependencies, or operational response, so Red Button recommends moving to production in a subsequent cycle.
Does Layer 7 HTTPS testing require our TLS private key?
No. Red Button connects as a client against your public certificate. The test does not require your TLS private key, and Red Button does not terminate or intercept your TLS. Some vectors deliberately vary the client-side TLS handshake or fingerprint, but no private key is ever shared.
Is DDoS simulation testing a paid service?
Yes. Red Button provides DDoS testing as a paid, expert-led engagement rather than a self-service platform, and scope drives the cost: the targets, the vectors in play, the environments covered, and whether retesting after remediation is included. Free and self-service tools have legitimate uses, and the comparison of testing options sets out what each one can and cannot establish.
The Short Version
Production DDoS testing provides evidence that staging cannot fully reproduce: the security configuration actually protecting the service, behavioural controls meeting real user traffic, the mitigation provider’s response, and the organisation’s own incident-response process.
That realism is also what creates the risk. The answer is not to make the test unrealistic. It is to control it – agree the scope and the traffic profile, coordinate with the providers, choose the right window, monitor from both sides, and preserve the customer’s ability to stop at any moment.
Red Button has worked on DDoS and nothing else since 2014, across more than 2,000 controlled simulations, and actively recommends production testing where the organisation is ready for it. Where it is not, staging is the first step – with production as the point at which DDoS resilience is ultimately validated. Talk to our team about what your environment can safely validate, or read how the DDoS Research Team approaches vector selection.
