ChatGPT and Claude can both write a strong training plan, but the quality depends almost entirely on what you tell them first. Peer-reviewed research tested this directly on ChatGPT and found that plans built from detailed athlete input scored significantly higher with coaching experts than plans built from a bare request. The model is not the variable. Your input is.
What AI is genuinely good at
Large language models are strong reasoners. Ask one to explain why your threshold session felt harder than the same session three weeks ago, or to restructure a week around a work trip, and you will get a considered answer that holds several variables at once. They are patient, available at 6am, and will happily explain the same concept four different ways until one lands.
What they are less reliable at is knowing which version of you is current. Assistant memory has improved a lot, and a new conversation no longer starts from nothing. What still slips is the part that changes: whether the PB you mentioned last month still stands, whether that race is still the target, whether January’s injury is still limiting you.
That is the actual failure mode: not stupidity, amnesia.
Athmex exists to fix exactly that: making sure the assistant is working from the current version of you, not last month's.
What the research says about AI training plans
A 2024 study in the Journal of Sports Science and Medicine tested this directly, using ChatGPT. Ten coaching experts, all with a postgraduate qualification in sports science and around seven years of endurance coaching experience, rated three ChatGPT-generated six-week running plans against 22 quality criteria on a five-point scale.
The three plans differed in one respect only: how much information the athlete supplied in the prompt. Median overall rating rose from 2 for the minimal-input plan, to 3 for the moderate one, to 4 for the detailed one. The detailed-input plan scored significantly higher than the minimal one on 15 of the 22 criteria.
Two caveats matter, and the authors raise both. The plans were generated in May 2023 on a model two generations old, so current results would likely be better. And agreement between the ten experts was weak, which limits how firmly any single rating can be read. The paper's own conclusion is cautious: it advises against using an AI-generated plan without an experienced coach reviewing it.
Read honestly, the finding is narrow and useful: output quality tracks input quality. Nobody has shown that AI coaching is safe unsupervised. What has been shown is that the difference between a mediocre AI plan and a decent one is mostly information you already have.
The five things to tell it before you ask anything
Most people open a chat and type "give me a 10K training plan". The assistant has no choice but to guess, and it guesses towards the average, which is nobody. Supply these five things and the answer changes completely.
1. Current benchmarks, with dates. Not "my 5K PB is 19:30" but "19:30, set in March this year". A two-year-old PB and a two-month-old PB imply completely different training paces, and the assistant cannot tell which it is looking at.
2. Your working threshold. The pace or power you actually train to, not an estimate from a race calculator. Everything in a plan is derived from this number, so an error here propagates through every session.
3. The race and the date. "Getting fitter" produces a generic plan. "Sub-40 10K on 6 September" produces a periodised one, because now there is a deadline and a required pace to work backwards from.
4. Real constraints. Hours available, not hours you wish you had. A plan built for eight hours a week is worse than useless to someone with five: it will be abandoned in week two and you will conclude the AI was wrong, when you simply gave it a fictional athlete to plan for.
5. Injury history and what triggers it. This is the one people leave out most often and it changes the most. "Left hamstring flares when weekly run volume goes above 50km" is a hard constraint that reshapes the entire structure. Without it you get a textbook build that walks you straight into the thing that always breaks.
What a good prompt actually looks like
Compare these two. The first is what most people send:
"Write me a 10K training plan."
The second is powered by your Athmex and produces something an experienced coach would recognise:
"I'm targeting sub-40 for a 10K on 6 September, so 3:59/km. Current 5K PB 19:30 from March, 10K PB 41:14. Threshold pace 4:05/km, max HR 192, threshold HR around 172. I'm 92kg and 191cm, so I carry more than most runners at this pace. Five to seven hours a week across swim, bike and run, I'm a triathlete, not a pure runner. Left hamstring gets irritable when run volume climbs, so I'd rather two hard run sessions than three. Build me the next four weeks and tell me which session matters most."
Same model, same question. The second version can produce paces derived from your actual threshold, a volume that fits your week, and a structure that respects the injury. The first cannot, no matter how capable the model.
What changes when the race is longer
The five inputs are the same for a 10K and for a marathon. What changes is which of them the plan is most sensitive to, and that is worth saying out loud, because an assistant asked to scale a plan by distance will usually scale the mileage rather than the thing the distance actually demands.
Over 5K and 10K, threshold pace does most of the work. The plan lives close to it, the sessions are short enough that fuelling rarely decides the outcome, and a recent parkrun or a 10K result is usually enough to set every pace in the block.
Over a half marathon, weekly volume starts to matter as much as pace, and the long run stops being a box to tick. Give it your current longest run alongside your PB. Twenty-one kilometres off a 12km long run is a different problem from the same race off 18km, and nothing in a PB tells it which one you are.
Over a marathon, the input that decides the plan is usually time rather than fitness. The long run has to progress somewhere, fuelling has to be practised rather than described, and the taper has to survive contact with your enthusiasm. Say how many weeks you have and how many of them are realistically uninterrupted. A marathon block is where an optimistic hours figure does the most damage, because you find out in week fourteen rather than week two.
None of this replaces the five things. It changes which of them you should be most precise about.
If you ride rather than run, the same logic holds with FTP in place of threshold pace, and the numbers a cycling plan is built on are covered separately.
Where AI coaching still falls short
It cannot see you run. Form, fatigue in your face, the way you favour one leg on a warm-up: a coach standing at the track picks up all of it and no chat interface can. It has no accountability relationship with you; nothing happens if you skip Thursday. And it is agreeable by default, so it will tend to validate a plan you have already decided on rather than tell you the plan is wrong.
It also cannot diagnose. If something hurts in a way that is new, sharp, or getting worse, that is a physiotherapist's job, not a chatbot's.
Does this replace a real coach?
For most amateur athletes it replaces nothing, because most amateur athletes do not have a coach. The realistic comparison is not AI versus a good coach: it is AI versus a generic plan downloaded from a website, or versus guessing. Against that baseline, an assistant that knows your actual numbers is a clear improvement.
If you already work with a coach, the useful role is different: understanding why a session is structured the way it is, or thinking through a scheduling problem between sessions. Not overriding them.
The part that breaks: doing it again next week
Writing that detailed prompt once is fine. Writing it before every conversation is not, and this is where people quietly give up. They start pasting screenshots, or they stop bothering and accept worse answers. Built-in memory takes some of that load, but it holds a synthesis the assistant curates rather than a record you control, so it may carry a 5K PB without knowing whether it still stands or what replaced it. What ChatGPT’s memory keeps, and what it loses goes through the current behaviour and OpenAI’s own numbers.
This is the problem Athmex exists to solve. You enter your thresholds, PBs, goals, races, injuries and equipment once as structured, dated facts, and your assistant reads them at the start of any conversation. Athmex holds only what you tell it, nothing is imported from another platform and nothing syncs in the background, and it gives no training advice itself. Claude or ChatGPT still does all the reasoning. It simply reasons from the athlete you are now, rather than the one you described in a chat six weeks ago.
If you also want your assistant to see recent sessions, your training platform's own connector can supply those directly. How the two-source model works goes through that split in more detail. A connector shows what you did; your context says who you are.