This is the number one cause of automations that work perfectly in testing and fail in production. Nothing changed in your scenario: volume crossed a threshold nobody mentioned.
Definition
A rate limit is the maximum number of calls a service accepts over a given period. Beyond it, further calls are refused.
It is expressed in several ways depending on the service: per second, per minute, per day, sometimes as a data volume or a Token count rather than a number of calls.
Refusal is not a failure, it is a normal response from a service asking you to slow down. The problem is what your automation does with it: if it ignores the refusal, it loses the items concerned without saying so.
The three ways to get refused
The burst
A thousand items arrive at once, your Workflow fires a thousand calls in ten seconds. The service accepts a hundred and refuses the rest. It is the most frequent case, and Batch processing solves it.
Silent accumulation
Three different automations call the same service. Each stays under the limit, their sum does not. Nobody anticipated it because nobody was looking at the whole.
Growth
Your volume doubles with nothing else changing. The threshold never reached becomes reachable one Tuesday morning, generally on a day you are not watching.
How to protect yourself
Read the limit before building. It is in the API documentation. Two minutes of reading avoids a rebuild.
Space calls deliberately. A one-second pause between calls is almost always enough, and costs nothing on an overnight job.
Retry intelligently. On refusal, wait then retry, doubling the wait after each failure. Retrying immediately makes things worse.
Monitor the refusal rate. A service starting to refuse 1 percent of calls warns you weeks before it refuses half.
The worst case is not loud failure, it is silent partial failure: three hundred items handled, seven hundred refused, and a status shown as completed. Always compare items handled against items received.
Frequently asked questions
Can a limit be raised?
Often, yes, by moving to a higher plan or asking support. AI providers generally raise limits as your consumption history builds.
How do you know where you stand?
Most services return your quota status in the headers of their responses. That information can be used to slow down automatically before hitting the wall.
Do AI models have the same limits?
They have two: calls per minute and tokens per minute. The second is often hit first when your instructions are long, which surprises people since the call count stays low.
How do you size an automation from the start?
By asking the maximum volume question when writing the Automation scenario. Our n8n course covers that sizing before building, not after the first outage.