Nobody has ever finished an FTP test and thought “well, that was a good use of a Tuesday.”
You know the drill. Twenty minutes of holding a number you picked before you started, in a room that smells like a gym bag, while a small voice asks whether you went out too hard at minute three.
Then you take whatever figure comes out the other end and let it define your training for the next three months.
It is a strange way to run a season.
The good news is that you probably do not need to do it any more. An AI cycling coach can work out your FTP from the riding you already do — and there is a reasonable argument that the number it gets is better than the one from the test.
Here is how that works, and where it still falls over.
First, what the test was actually for
Functional Threshold Power is meant to be the highest power you can hold for roughly an hour. It matters because almost every training zone is calculated as a percentage of it. Get FTP wrong and every session after it is wrong too — slightly too easy, slightly too hard, quietly pointless.
The 20-minute test exists because riding flat out for a full hour is miserable and most people will not do it twice. So the sport settled on a shorter effort and a fudge factor.
That is the whole idea. A hard effort, a bit of arithmetic, a number.
It works. It is also a single sample.
The problem with a single sample
Your FTP test measures one day.
Specifically, it measures the day you decided to do it. Which is usually a day you slept badly, or a day you had a big week, or a day you were quietly dreading it and went out conservatively because failing publicly is worse than scoring low.
One rider put it perfectly in a conversation with our coach:
“That’s quite low. I know for a fact I can ride at 220 watts for hours.”
The app said 180. His legs said 220. He was not confused about his own body. He was confused about why the software disagreed with it.
And here is the thing nobody says out loud: he was probably right. He had months of evidence — every long ride, every group day, every climb — and the software was ignoring all of it in favour of twenty minutes from one Tuesday in the rain.
Once that number is in the system, everything downstream inherits it. Your intervals are set from it. Your endurance ceiling is set from it. Your progress is measured against it.
A test that lands 20 per cent low does not just give you an annoying number. It gives you three months of training that is too easy to change anything, and you can feel it, which is why you keep quietly overriding the workouts and then feeling like you have cheated.
What an AI cycling coach does instead
The short version: it stops asking you to prove something you have already proved.
You produce threshold-relevant data constantly without ever calling it a test.
Your hard group ride. The forty minutes where the pace lifted and you held on. That is a maximal effort with a real number attached, performed under more honest conditions than a scheduled test — because you were not pacing to a target, you were trying not to get dropped.
Your climbs. A twenty-minute climb ridden properly is an FTP test that you enjoyed. Sort of.
Your best sustained efforts across weeks, not days. This is the real advantage. Rather than one sample from one day, the model has dozens of samples across every kind of condition — fresh, tired, hot, cold, motivated, hungover, indoors, outdoors.
From those, it fits a curve. Your power over different durations follows a fairly predictable shape, and the point on that curve at around an hour is your threshold. You do not have to ride the hour. The shorter efforts constrain where the curve has to sit.
That is not a trick. It is the same maths a coach does in their head when they look at your file and say “you’re probably around 250, not 230.”
Why this tends to be more accurate, not less
It sounds like the estimate should be worse. It is a guess from indirect evidence, versus a direct measurement.
Except the direct measurement has all the problems above, and the estimate has three things going for it.
More data beats better data. Thirty efforts across six weeks tells you more about your actual threshold than one effort on one day, in the same way that thirty rides tells you more about your fitness than your single best ride.
It updates continuously. Your fitness does not politely wait for your next scheduled test. It moves week to week. A number that updates as you ride is closer to true more of the time than a number that is correct on test day and drifts for the eleven weeks afterwards.
It cannot be sandbagged. Not deliberately, anyway. Everyone paces their first FTP test badly. Plenty of people pace their fifth one badly too, because the test rewards a very specific skill — even effort distribution under discomfort — that has only a loose relationship with your actual threshold.
Where it still falls over
This is where most articles about AI coaching stop being useful, so let us not.
It needs something hard to look at. If every ride you do is a flat, easy, ninety-minute spin, there is nothing in the data for the model to pull a threshold from. It is not magic. It reads efforts, and if you never make one, it has nothing to read. In practice this means an AI coach is less accurate for a rider doing pure base than for one who races or does group rides.
It needs a few weeks. A brand new account with three rides in it is guessing. That is a real limitation and any coach that pretends otherwise is overselling.
Power meters drift. Estimation from your data inherits whatever your data is wrong about. If your meter reads 5 per cent high, so does your estimated FTP. Calibrate the thing.
And it has to explain itself. This is the big one, and most of the category gets it wrong.
The bit that actually breaks trust
There is a thread that comes up constantly on cycling forums, in almost the same words each time:
“With the new update the AI dropped my FTP from 204 to 176 and after that my FTP almost no increase at all… this is so confusing.”
Read that again, because the important word is the last one.
Not “wrong.” Not “unfair.” Confusing.
The complaint is almost never that the number is hard to accept. Riders are realistic about their own fitness. The complaint is that the number moved and nobody said why.
An estimated FTP that arrives with no explanation is worse than a test result, even if it is more accurate — because at least you were there for the test. You know what happened. You know you went out too hard. The number might be wrong, but it is yours.
So the standard for any AI coach that does this should be:
- Tell you what your threshold is.
- Tell you which rides it read to get there.
- Tell you when it changes, and what changed it.
- Let you override it without making you feel like you have broken something.
That last one matters more than it sounds. Riders override their plans constantly and then feel guilty about it, as though they are cheating a system they are paying for. They are not. They are supplying information the software does not have.
So should you ever test again?
Sometimes. Two cases.
You are starting from nothing. No power history, new meter, coming back after a long layoff. One hard effort gives the model somewhere to start rather than guessing from an empty file.
You actually want to know. Some riders like testing. It is a benchmark, it is satisfying, and doing it twice a year to check the estimate against reality is a genuinely reasonable thing to do. As one rider we spoke to put it: “it’s all hope until you test.”
He is not wrong. But note what he is describing — a confirmation, not a foundation. Testing to check the number is different from testing because the number cannot exist without it.
What this is really about
The FTP test is not the point. Nobody took up cycling because they wanted an accurate threshold estimate.
The point is the thing underneath it: knowing you are doing the right work, and that it is moving you forward.
A number you got from one bad Tuesday does not give you that. Neither does a number that changes overnight with no explanation. What gives you that is a threshold that reflects how you have actually been riding, that updates as you improve, and that tells you where it came from.
The test was never the goal. It was just the only tool we had.
Streeka reads your threshold from the riding you already do, tells you which sessions it used, and explains it when it changes. No test required — though you are welcome to do one if you enjoy that sort of thing. See how it works.