Reinforcement and reward

Variable reward schedule

Rewarding a behaviour inconsistently rather than every time, which tends to make that behaviour more resistant to fading.

What is a variable reward schedule?

A variable reward schedule means rewarding a behaviour inconsistently rather than every single time it happens. Instead of a treat for every sit, the dog might get one after the first sit, then not the next two, then again on the fourth.

This unpredictability tends to make a behaviour more resistant to fading, because the dog can't predict which repetition will pay off, so they keep offering the behaviour in hope of reward.

How it's used in training

Early in training a new behaviour, it's normal to reward every single repetition. This is called a continuous reinforcement schedule, and it helps the dog learn the behaviour quickly.

Once the behaviour is reliable, the schedule is gradually thinned out. A handler might move from rewarding every sit, to every other sit, to an unpredictable pattern where sometimes two sits in a row get rewarded and sometimes five don't. Over time, this builds a behaviour that holds up even without a treat visible.

Common mistakes

Thinning the schedule too early, before the behaviour is solid, can cause it to fall apart. It's best to keep rewarding every repetition until the dog is performing the behaviour confidently and consistently, before introducing any unpredictability.

Frequently asked questions

Why not just reward every time forever?

A dog rewarded every single time can become dependent on seeing a treat before responding. A variable schedule keeps the behaviour strong without needing a reward for every repetition.

When should I start using a variable schedule?

Once the dog is reliably performing the behaviour on cue in a range of situations. Moving to a variable schedule too early, before the behaviour is solid, risks weakening it instead.

Is a variable reward schedule the same as randomly rewarding?

It's similar in that the pattern is unpredictable, but it should still be applied thoughtfully, gradually increasing the gaps between rewards rather than being entirely random from the start.