Notes
How to choose a reCAPTCHA v3 score threshold for your checkout
Google says the score is a risk signal, not a verdict. Here is how to read reCAPTCHA v3 scores on your own traffic and pick a threshold that does not cost you real orders.
7 min read
You have added reCAPTCHA v3 to your checkout, and it now returns a number between 0.0 and 1.0 for every order attempt. Choosing a reCAPTCHA v3 score threshold for that number is genuinely uncomfortable. One end of the dial loses you money to fraud, the other loses you money to customers who gave up and went elsewhere. You will not find out quickly if you choose badly, because blocked bots are silent and blocked customers are silent too. They simply leave.
What the 0.0 to 1.0 score actually represents
Google's reCAPTCHA v3 documentation puts it plainly: "reCAPTCHA v3 returns a score (1.0 is very likely a good interaction, 0.0 is very likely a bot)". The Google Cloud page on interpreting assessments expands on that. "The score 1.0 indicates that the interaction poses low risk and is very likely legitimate, whereas 0.0 indicates that the interaction poses high risk and might be fraudulent."
Read the hedging in Google's own wording. Very likely. Might be. That is the language of probability, not identification. It does not tell you a person is a fraudster, only how much this request resembles traffic Google has learned to distrust.
A second detail changes how you should tune. The same Cloud page says "reCAPTCHA has 11 levels for scores with values ranging from 0.0 to 1.0". Then it adds a caveat: "Out of the 11 levels, only the following four score levels are available before triggering an automatic security review by adding a billing account to your project: 0.1, 0.3, 0.7, and 0.9." Read that scope carefully. Google wrote it for Cloud projects without a billing account. It publishes no equivalent figure for the classic free v3 keys you get from google.com/recaptcha, which is the kind Checkout Bouncer uses. Treat it as a reason to go and count your own values, not as a fact about your checkout.
If your scores only ever land on a handful of values, thresholds of 0.4, 0.5 and 0.6 behave identically. Each of them splits 0.3 from 0.7. Agonising over the second decimal place is imaginary precision. You are choosing which bucket falls on which side of the line.
The scoring asks nothing of the shopper. Google's guidance on choosing a key type says "Score-based keys let you verify whether an interaction is legitimate without any user interaction". It lists "payment-related transactions that prefer less friction for better conversion rates" among the cases such keys suit.
Why Google suggests 0.5, and why that is only a starting point
The v3 documentation gives you a reCAPTCHA v3 score threshold to begin with: "By default, you can use a threshold of 0.5."
That is a default in the engineering sense, not a reCAPTCHA v3 score threshold chosen for your store. It is reasonable when you know nothing, and it should give way once you know something. Google's best practices for automated threats are blunt: "Your score thresholds might vary depending on your users and attackers. We recommend using the reCAPTCHA analytics dashboard to determine the best scores to take action on."
Your shop differs from the next one by design. Google says "reCAPTCHA learns by seeing real traffic on your site." The distribution you see is a product of your traffic and your checkout flow. Someone else's threshold is a fact about their shop, not advice about yours.
A score is a signal, not a decision
A reCAPTCHA v3 score is a risk signal, not a verdict about a customer. Google's guidance points repeatedly at graded responses rather than a single gate. The v3 docs advise taking action behind the scenes instead of blocking traffic. The Cloud documentation repeats it: "To protect your site better, we recommend that you take the action in the background instead of blocking traffic."
The verification response also gives you more than the number. Alongside the score it returns the action name, a challenge timestamp and the hostname. Google is explicit that you should use the action: "when you verify the reCAPTCHA response, you should verify that the action name is the name you expect." Naming actions also buys you "Adaptive risk analysis based on the context of the action, because abusive behavior can vary".
Graded responses carry a tactical advantage. Google's advice on carding says: "When possible, allow the transaction to proceed at time of purchase, but cancel the transaction later to avoid tipping off the attacker." A hard block is instant feedback for whoever is probing you, and it tells them which technique to change.
Set your reCAPTCHA v3 score threshold from your own traffic
Google gives the same instruction in both documentation sets. The v3 page says "you can first run reCAPTCHA without taking action and then decide on thresholds by looking at your traffic in the admin console." The Cloud page repeats it for score-based keys: "you can first run reCAPTCHA without taking action and then decide on thresholds by looking at the traffic."
There is also a warm-up period, and it catches people out. The v3 docs warn that "scores in a staging environment or soon after implementing may differ from production". The Cloud documentation attaches a figure: "reCAPTCHA learns by monitoring real traffic on your site. Therefore, scores in a staging environment and within 7 days of implementation might differ from the long-term production scores."
A reCAPTCHA v3 score threshold chosen on your first afternoon of data is chosen on the wrong data. Install, record the scores, block nothing, and wait out the warm-up. Then line the scores up against orders that were paid and fulfilled without trouble. Those are your known-good population, and their lowest scores are the real constraint. If a fair share of them sit in a low bucket, a threshold above it would have turned genuine customers away.
What happens at the extremes
Push the reCAPTCHA v3 score threshold high and you are asking reCAPTCHA to be confident about everybody. Count the distinct values in your own log before deciding what a step costs. Where scoring turns out to be coarse, moving up a bucket is not a small tightening. It is a large one, and the people it removes never tell you.
Push it low and you refuse only traffic Google is close to certain about. For a checkout that is defensible, since a wrongly refused order is lost immediately while a suspicious one can often be held or cancelled later. The trade is that plenty of automated attempts will pass and need catching some other way.
One trap near the extremes looks like a scoring problem and is not. Google's verification documentation states that "Each reCAPTCHA user response token is valid for two minutes" and that it "can only be verified once to prevent replay attacks". The timeout-or-duplicate error means "The response is no longer valid: either is too old or has been used previously." A shopper who lingers over the card fields, or who corrects a declined payment and submits again, can trip that without ever being scored. Treat every verification failure as a bot and you block people your threshold never judged.
A workable order of operations
- Install, log the score for every checkout, enforce nothing.
- Wait past the warm-up window before reading anything into the numbers.
- Store each score next to the eventual outcome of that order.
- Check the action name and hostname on every verification, as Google instructs.
- Handle expired and reused tokens separately from low scores.
- Set the threshold below the floor of your known-good orders, not at a number you read somewhere.
- Prefer holding or flagging over refusing outright, wherever the business can absorb it.
None of this requires a plugin. A reCAPTCHA v3 score threshold is a decision, not a feature. It requires the score recorded against every checkout, and the discipline to look before you act.
Checkout Bouncer is our free WooCommerce plugin, GPL and published on WordPress.org. It scores checkout submissions with Google reCAPTCHA v3, the only engine it uses. It records the score, verdict, surface and reason for each one, which is what makes the reading above possible. One thing changes how you run the first step. Scoring enforces as soon as you enable it, so observing before enforcing means reading the Logs tab rather than running unenforced.
It covers the WooCommerce checkout and nothing else. It will not help with logins, registration, comments or contact forms, and it is neither a firewall nor a malware scanner. It cannot tell you your reCAPTCHA v3 score threshold either. That number comes from your own traffic, which is rather the point of everything above. The setup documentation covers where the threshold lives and what the fail-open and fail-closed settings do. The features page lists which checkout surfaces get scored at all.