Scaled Down Overnight, Trampled at Dawn — JavaScript Bug Hunt

Modelled on Slack's outage of 4 January 2021: the fleet scaled down over the quiet holidays, then the first working day of the year brought everyone back at…

  • Language: JavaScript
  • Layer: Backend
  • Difficulty: Medium
  • Concepts: Autoscaling, Thundering Herd
  • Modelled on: Slack · 2021
  • Visible tests: scaling up is immediate; scaling down is gradual; the floor is respected
  • Reward: 50 XP for a complete fix

Briefing

Modelled on Slack's outage of 4 January 2021: the fleet scaled down over the quiet holidays, then the first working day of the year brought everyone back at once. The autoscaler could not add capacity fast enough, and the resulting overload took the service down for hours.

autoscaler.js sizes the fleet purely on current load, with no floor and no rate limit.

Fix scaleTarget so it never drops below a floor and never scales down faster than the configured step.

Bug report

BUG-SLACK0104 · Priority: High · Reported by: capacity SRE

scaleTarget(current, load, config) must return the next fleet size, where:

  • the raw target is load / config.perInstanceCapacity, rounded up
  • the result is never below config.minInstances
  • a scale-DOWN never removes more than config.maxScaleDownStep instances at once
  • scaling up is not rate limited

Observed: an overnight lull takes the fleet from 400 instances to 2 in one step, and the morning ramp cannot recover.

Logs

[autoscale] target 2 (from 400) load=95 rps
[autoscale] 09:02 load=42000 rps, healthy instances 4, queue depth 1.2M

The code as shipped

src/scale/autoscaler.js (editable)

// Computes the next fleet size from the current load.
exports.scaleTarget = function (current, load, config) {
  return Math.ceil(load / config.perInstanceCapacity);
};

Read-only context: src/scale/CONFIG.js.

Open the hunt to edit the files, run the visible tests and submit against the hidden ones. More JavaScript bug hunts.