Ready. Set. Send. Boom! 500 employees received an email promising a $650 holiday bonus. You see it and click it, but instead of a bonus you get cybersecurity training. What a letdown. This was the strategy GoDaddy used with its employees during the Holiday season. Fake Holiday bonuses. Replacing those bonuses with additional security training that needed to be completed before the end of the year.
If you’ve seen me write on phishing before, you know that I absolutely despise how they’re run today. How click-rate is an absolute vanity metric that has no correlation to the security of the company. Nor is it a key performance indicator (KPI) that relays how well an employee is retaining security information. Let alone one that should be used in performance reviews. I’m not sure when, but there was a moment when phishing simulations stopped being about awareness. This wasn’t always the case; it wasn’t the original design. However, organizations now use it as a standard to measure the level of human risk defense. I’m here to tell you, there is a better way! It should be impossible to fail a phishing simulation. Ironically using failure rate as a metric may be one the main causes of higher click rates.
In this blog, we’ll dive into how organizations use phishing simulations to build a false sense of security and how applying behavioral science and AI can lead to more robust behavior change. Without having to fail a single user.
How phishing simulations lost their original purpose
Before diving into how to improve today’s phishing simulations, we first have to look at how we got here. The word simulation derives from the Latin word Simulātiōnem, which means to imitate, copy, or represent. The original intent of a simulation was to provide the training or reinforcement the individual needed to happen. Therefore, the simulation was run; failures would be tracked; instructors would collect the data; and then use that data to inform better training and future sessions.
When phishing attacks arose in the mid 2000’s, primarily with the rise of email adoption, organizations began using the same training principles, with simulations, to combat the rising tide of attacks. The original intent was “build a safe space for employees to see phishing emails so they’d recognize them in the real world.” As simulations grew in popularity, they gradually became the measure of the success rate of a general security awareness program. In essence, it helped answer the question, “How effective is your security training?” This became a more consistent question as auditors for SOC2, ISO, HIPAA, and other common compliance frameworks use security training as a measure of robustness. This formed the dreaded “click rate” KPI.
Once “click-rate” became a KPI, phishing simulations became the target of employee shaming. Ironically, this development was actually the beginning of the end in terms of the effectiveness of security awareness programs. Sadly, this number has become a weaponized metric to ridicule employees and “scare them straight”. 42% of organizations use “click rate” to take disciplinary action against employees. With 15% saying they are openly sharing the names of people for failing. A Fortune 50 CISO even admitted to using “click-rate” as a metric to fire people. They followed a “three-strikes, and you’re out” policy. In the case of our earlier example with GoDaddy, it meant losing a “bonus” and being forced to do 40 minutes of training over the Holiday break. Thus, “click-rate” inadvertently created a negative connection between how employees and cybersecurity.
Why does punishing phishing failures make security worse?
The negative association between phishing simulations and employees is pervasive across organizations. Roughly 90% of users say their primary interaction with a security team is either when they click a phishing simulation link or when they are put into remedial training for a failure. Think about what that means. The only time most employees hear from the people responsible for protecting the company is when they've done something "wrong." Security shows up as the department that catches you, not the department that has your back.
Decades of behavioral research show that negative reinforcement is one of the hardest ways to get the best results. Negative feedback or shaming in isolation leads to constant pressure. I personally used to feel this when I was in SAT prep class in High School. We took a practice SAT every week, and our scores would be posted on a wall. If your grade was on the bottom, everyone knew. Theoretically, it would motivate you to work harder. However, this was something that was the opposite. The more I personally tried, the less I improved. The less I improved, the more I would just be nervous and fail. Ultimately, my perception of SAT prep is all negative. The test itself and the process.
Phishing simulations are no different. The data supports that failing phishing simulations that directly lead to negative feedback leads to heightened anxiety, lower productivity, more resentment, and less willingness to report. CybSafe's lab experiments found this impact was "highly detrimental", and their survey work found that 42% of organizations take disciplinary action against employees for cybersecurity errors anyway, from public shaming (15%) to revoked access (33%) to looping in their manager (63%). That's almost half the industry actively building a system that punishes the behavior they're trying to improve.
Here's what almost nobody in the space talks about: forcing people into mandatory training after they click doesn't even fix the clicking. A peer-reviewed study at a major US healthcare system tracked what happened when repeat clickers were funneled into a required training program. Click rates for "offenders" stayed between 10% and 25% after the mandatory training. The program didn't change the behavior because the behavior wasn't a knowledge problem in the first place. It's a workload problem. An attention problem. A context problem. No amount of remedial training solves any of those. It just tells the employee, "You failed, now sit here for 20 minutes and think about what you did." It's the corporate version of writing lines on a chalkboard.
It gets worse. ETH Zurich ran one of the largest real-world studies ever done on phishing simulations and found that employees who received contextual training after clicking actually got worse… click rates went up 16%, and dangerous-action rates went up 27%. A separate study out of UC San Diego tracking 19,500 employees found that each additional training session was associated with an 18.5% increase in failure likelihood. The industry's answer to clicking has always been "more training," and the best evidence we have says more training makes things worse. Ultimately, if your phishing program punishes people, it's not a security program. It's a liability program, you're creating evidence for termination while quietly destroying the reporting culture that can be an extension of your security team.
What does behavioral science say about phishing simulations?
Now, with a problem comes the opportunity to evolve phishing simulations from vanity metrics to actual behavior change. News flash, it’s not sticking people with remedial training, and it’s not gamifying everything. It uses actual behavioral psychology to maximize behavioral change.
In behavioral science, there’s a theory called Gagné's Hierarchies of Learning, which presents that there are different stages of learning that work in a hierarchical sequence. Think of this as crawl, walk, run. One of the key components of Gagne’s theory was the application of learning in two phases: instruction (designed to produce learning) and assessment (designed to measure whether learning occurred). This is ironically how the industry has been developing phishing simulations for the last 30 years. You have instructions, “don’t click it” relayed to the assessment, “click rate”. However, Gagné explicitly distinguished instruction from assessment. The current phishing simulation model collapses these into a single event, which is exactly the category error Gagné spent his career warning against.
In Gagné's framework, instruction and assessment have distinct conditions, purposes, feedback loops, and success criteria. An instructional event that fails (e.g., the learner clicks a link) is data for the instructor; it tells you to adjust rather than how to punish. An assessment event that fails is data about readiness (how prepared is the organization for a real phishing attempt). Confusing them means you punish people for the act of learning, which is pedagogically incoherent.
Here's where Gagné's hierarchical task analysis gets genuinely useful. "Don't click phishing emails" is not an atomic skill. It's a higher-order rule that sits atop a stack of prerequisites. If you wanted to build the actual learning hierarchy, it might look something like:
- At the bottom: discriminations: telling a legitimate sender domain from a spoofed one, recognizing when a URL preview doesn't match its anchor text, noticing when an email's visual branding is slightly off.
- Above that: concrete concepts: "this is a spoofed domain," "this is a display-name mismatch," "this is a suspicious attachment type."
- Above that: defined concepts: "this is a pretext," "this is social engineering," "this is credential harvesting," "this is business email compromise."
- Above that: rules: "if a request creates urgency AND bypasses normal process, verify through a second channel"; "if an email asks for credentials, never enter them via the link."
- At the top: higher-order rules / problem-solving: novel situations where the specific pattern hasn't been seen before but the learner synthesizes from prior rules to recognize something is off.
The current phishing simulation model is flat. It tests only the top of this hierarchy, “did you apply the higher-order rule correctly under time pressure?” without having taught or assessed any of the prerequisites. That's like giving someone a calculus exam to find out whether they know algebra. A click isn't one data point; it's a failure somewhere in a stack, and the current model can't tell you where. A further example is, let’s say a finance employee who clicks an invoice-fraud lure might have failed at the discrimination level (didn't notice the domain), the concept level (doesn't know what invoice fraud is), or the rule level (knows the pattern but didn't follow procedure under pressure). These are three completely different training needs, and the flat model treats them identically. In order to change the effectiveness of phishing simulations in your organization, you must start by fundamentally changing how they’re executed.
A two-tier model for phishing simulations
To start, we recommend building phishing campaigns in two tiers:
Learning simulations. Low-stakes. Possibly pre-announced. Designed to teach specific patterns: invoice fraud for finance, CEO impersonation for execs, credential harvesting for everyone. The goal isn't to catch anyone. It's to build pattern recognition. A click here is a data point saying "this pattern needs more exposure," not a mark on a record.
Testing simulations. Higher stakes in the sense that they measure readiness, but still not punitive. These assess whether the learning is stuck at the department or organizational level, not the individual level (this echoes the "test teams, not individuals" principle). Results inform training investment, not HR files.
To build any of these simulations, you need a lot of information. The goal being that executing one phase of simulations can lead to an additional phase of simulations. Data from one interpreting data from the next. With this in mind, security teams need to produce:
- More simulation templates and landing pages
- Customized phishing scenarios for each type of persona
- Understanding of the latest phishing attempts across the industry
- Easy ways to track all of the data that comes from clicking, reporting, and even opening
Building an entire phishing program in 10 minutes with Herd
I know what you might be thinking. "Great. So now, in order to get phishing simulations to be valuable, I have to put a ton of extra work into building hierarchies, varying severity, tracking report-rate, designing learning paths vs. testing paths, and rethinking the whole program from scratch." Fair concern. Because honestly, that is what the research points to. A humane, tiered, non-punitive program isn't just an attitude shift. It's a design shift, and design shifts cost time most security teams don't have.
This is where the honest truth about tooling comes in. Existing phishing simulation platforms were built for the flat mode. One-size-fits-all campaigns. Generic templates. A single click-rate dashboard. A handful of role-based variations, if you're lucky. They make it easy to run the program we've been arguing against and hard to run anything else.
This is the exact gap we built Herd to close. Not by slapping AI on top of the old model, but by making the tiered, learning-first, report-rate-focused model approach with no extra overhead. You bring in your organizational and industry context, and Herd helps you generate phishing simulations and landing pages across varying severity levels, different interaction types, and different learning objectives.
What that looks like in practice: teams are building 10 phishing templates in 20 minutes. They're running tiered campaigns without expanding headcount. They're tracking report-rate and interaction depth alongside click-rate, so the program actually reflects security behavior instead of punishing the one behavior that's easiest to measure.
Everything we just talked about, the reporting culture, the learning-vs-testing hierarchy, the move away from "click it or ticket", none of it has to be a massive lift. That's the whole point. You shouldn't need to choose between a humane program and a realistic workload.
You can see it for yourself with a free trial.
Stop the click!
A click isn't a failure. It's a signal. A learning moment. A data point showing where training hasn't stuck or where an attack pattern is genuinely convincing. Failure only enters the picture when an organization chooses to treat it as failure, and that choice isn't just cruel, it's counterproductive. The only way to truly fail a phishing simulation is to build a culture that punishes people for taking one.




