A Different Way to Look At It

By Chad Mackin

Dick Russell used to say, “Everything we know about dog training can be summed up in two sentences: If just as or just after a dog does something, something good happens, the dog’s gonna do that thing again. If just as or just after a dog does something, something bad happens, the dog’s not gonna do that thing again.”

You may recognize this as a short description of reinforcers and punishers. And that makes sense. Most dog trainers are trained to think in those terms first. And the quadrants are certainly valuable. You can explain and solve a lot with a solid understanding of how they work. But thinking exclusively or even mostly in terms of quadrants comes with some serious limitations. The old mantra of the behaviorists was “Consequences drive behavior” which is absolutely true. But it’s incomplete. Neuroscience has updated that to “Emotions drive behavior.” Reinforcement and punishment work because they change how the dog feels.

Reinforcers make a choice more appealing, while punishers work in the opposite way making a choice less appealing. They are only effective if they change how the dog perceives the value of a choice. Like all organisms, the dog is trying to find the best, most rewarding experience.

If you think about it, every behavior comes down to a cost benefit analysis. For example, I never saw my grandfather walk past a payphone without checking the coin return for some forgotten change. I also never saw him walk across the street to check a phone for loose change. If the payphone was in his path, it was worth the effort to just reach out and check the coin return. It might even be worth a couple of steps out of the way. But it wasn’t worth going very far out of his way. The cost was too much to be supported by the potential reward.

On the surface, this seems like a pretty straightforward choice, but if you look a little deeper, it becomes more complex. How far he was willing to go out of his way would be affected by how tired he was, whether he was sore, or running late for an appointment. Maybe if there was a bank of six phones in a busy location like an airport, he might figure the odds of finding more than a few cents were larger, making going out of his way a better deal.

While it matters quite a bit whether he’s been reinforced previously, there are a ton of other factors that probably weigh heavier in the decision making process. The quadrants, as we tend to apply them, can’t explain why one day he will walk 5 feet out of his way, but on another day he wouldn’t. If the rate of reinforcement hasn’t changed, his behavior should be pretty stable.

I’m not saying anyone would predict it would be stable. Everyone reading this has too much experience being human to expect that. I’m saying the way we apply learning theory doesn’t deal with this entirely predictable result. That may not be a weakness in the model, but it is a weakness in the way we think about it.

I’d like to shift that thinking by saying explicitly something we all intuitively understand.

There are two ways we can affect future choices. We can change the cost of the behavior, or we can change the payout.

Or more precisely, we can change the expected cost or the expected payout. Because what we are really talking about is perception. There are no absolute values when it comes to behavior.

Every choice costs something in terms of things like effort, energy, unpleasantness, strained relationships, potential disappointment, discomfort, or frustration, to name a few. Those costs are balanced, usually intuitively, by the level of reward expected or hoped for. The question the dog is always asking, often unconsciously, is “Is it worth the effort?” Or more accurately, “Will I probably be better off if I do this, or not?”

Reinforcement works because it offsets the cost of the behavior. The harder the behavior is, the higher the cost, and the more reinforcement it’s going to take to make the behavior feel worth it. And the reverse is true too. The more rewarding the behavior appears to the dog, the more it will have to cost to persuade the dog to skip it.

The secret is: No dog willingly takes the bad bargain.

But a whole lot of what makes a bargain look bad or good depends on what’s going on for the dog the moment before the options are presented. That’s not a secret either. Everybody has experienced moments when something normally enjoyable seems entirely uninteresting because they “just aren’t in the mood.” Or they’ve experienced moments where they are suddenly craving something they normally aren’t interested in at all.

What’s perplexing is that despite knowing that, most training interventions are focused not on the moments before the choice, but on the moments after it.

Punishment and Reinforcement work downstream of behavior in hopes of changing outcomes the dog predicts the next time the dog finds themself in a similar situation. That’s a valuable training strategy. The effects of downstream interventions are among some of the most thoroughly documented things in behavioral science. And one could make the argument that downstream of the last rep is still upstream of the next one. That’s an accurate framing for what it’s worth. But downstream interventions still require the dog to make a choice we can act upon. You cannot shape a behavior for a dog who doesn’t do anything. You cannot punish a dog until the dog does something wrong. In that sense, both choices are reactive not proactive.

Working upstream is often more valuable. Most people know it’s a bad idea to go to the grocery store hungry if they want to avoid overspending or buying junk food. But the effect is actually bigger than that. Being hungry makes you more likely to buy anything. Hunger causes people to want to acquire and hold onto things, because it’s not just about food. It’s about acquiring resources. The dog’s hunger makes him more likely to work for a toy, not just food.

Whether we are working upstream or downstream, success means changing the cost/payout value in the dog’s head.

Framing things in terms of a competition between these two evaluations gives us an approach that is more complete, nuanced and intuitive than relying entirely on reinforcement histories and the quadrants. It doesn’t nullify them, it puts them soundly into a larger framework that can help explain why a reinforcer that works today might not work tomorrow.

It also makes clear that changing a single behavior involves, whether we are aware of it or not, several factors that we can’t always account for. The dog’s emotional state, biological resources, relationship to the handler, physical health, emotional regulation, genetics and history all influence what the dog will do in that moment. They all affect how the dog evaluates what the behavior is going to cost, and what will be gained.

This framing also allows us to rethink the idea of punishment. Specifically, the notion that punishment relies on fear and pain to be effective. When many people think about using punishment they assume that if something is unpleasant enough to stop the dog from doing something they really enjoy, it must cause a significant emotional swing. And that swing can really only be explained by pain, fear, or both. Which is one reason trainers who avoid punishment tend to treat all punishment as very dangerous, and even might equate mild punishment with abuse. They are thinking of it as the entire intervention instead of just one piece of the dog’s evaluation.

But this framework allows us to see punishment in a minor supporting role, and points to a path to solve or reduce many behavior problems without punishment.

We can shift the expected value for the dog via humane and efficient means. And in many cases we can do that entirely with upstream interventions, which means in those cases there’s no need for punishment. We can provide other outlets for the drives pushing the unwanted behavior. We can build emotional stability and confidence. We can make the dog more optimistic. We can also condition the dog to believe the unwanted behavior will never produce a payout. All of these things can lower the value of the unwanted behavior making it less likely to occur. However, in a strict operant model, punishment is the only way to stop a behavior. That’s encoded in its very definition.

And that brings us back to the pain and fear idea. If we get to the point where we’ve convinced the dog via upstream interventions that the payout isn’t very high, and by interfering with rewards downstream that it’s unlikely to get even that meager reward, the dog still says “You know, I still think it’s a good idea!” We may need to only shift that ledger a tiny amount for the dog to decide “Ok, I don’t think it’s worth it anymore!”

At that point, punishment doesn’t need a huge emotional shift to be effective.

Sometimes it’s a matter of just one more minor inconvenience pushing the balance to the other side of the ledger.

By thinking in terms of changing the payout, or the cost, we immediately see solutions that go beyond reinforcers and punishers. We can also see the value of less intense reinforcers and punishers. We don’t need to be looking for huge emotional shifts when minor adjustments will do the trick.

And minor emotional shifts are all relatively safe. Frustration, disappointment, and avoidance are biologically safe emotional states as long as they exist within certain ranges. Dramatic emotional swings produce dramatic behaviors, but also produce a greater risk of superstitious learning, which in the case of punishment is a huge source of fallout. But in an emotionally stable dog, a minor aversive event is extremely unlikely to produce trauma. Dogs experience disappointment and frustration every day, and no one worries about fallout in a healthy dog when life throws them an obstacle.

Within the cost/payout framing, the quadrants actually gain clarity and nuance at the same time, and they are no longer thought of as isolated events, but as an integral piece of a more balanced, no pun intended, picture of behavior. Suddenly, what quadrant we are using becomes less important than what function is being served and the emotional impact.

When we embrace the cost/payout model, the distinction between balanced and force free seems less important, because many of the assumptions made about the so-called “negative quadrants” become less obvious, while also giving a clear path for balanced trainers to reduce the frequency and intensity of those quadrants without sacrificing results.

The cost/payout model isn’t a substitute for the quadrants, nor is it an indictment of them. It provides a backdrop that makes them more valuable by providing a clear path to making smaller interventions more effective.

More clarity, more options, and less to fight about.

Author Bio:

Chad Mackin is a professional trainer with over 33 years of experience. He has taught workshops on dog training all over the United States and Canada as well as England, Australian and New Zealand. He co-hosted Dog Training Conversations Podcast with Jay Jack, and currently hosts the Something To Bark About Podcast. He will also be a Presenter at the IACP’s 2026 Conference.

Also in this issue: