Part of the complete guide: How to improve your archery score

On the morning of the eliminations the bracket goes up at the entrance to the field, and next to the first name there is a 1. Everyone reads that number the same way, as a verdict already delivered: that archer is the strongest in the field, shot better than anyone else the day before, and from here on every result will be weighed against the label. Go out in the quarterfinals and you have disappointed. Reach the semifinal from seed 32 and you have surprised.

That number says a great deal less than it is asked to say, and the distance between the two can be measured. In a top-level field the top qualifier really is the best archer present a little over once in seven. The top eight really are the best eight four times in ten thousand. And every archer enters the bracket carrying a label that is wrong, on average, by eight and a half places. None of this comes from miscounting or from inattentive judges. It comes from the relationship between how fine the question is and how precise the instrument answering it can be.

The figures below come from work I published in the Lab, where the calculation is set out in full and every assumption is declared one by one.

How a competition is built

The structure is worth describing properly, because people who do not spend time on shooting fields know it only at second hand and people who do take it too much for granted. It opens with a qualification round in which everyone shoots 72 arrows at 70 metres, on a face 122 cm across divided into ten scoring rings of equal width, 6.1 cm each, with an X ring of 3 cm radius at the centre that is counted separately and exists to break ties. The theoretical maximum is 720 points, and the totals from that day produce a ranking from 1 to 64.

The ranking awards nothing and rewards nobody. What it does is set the pairings: first meets sixty-fourth, second meets sixty-third, and so on across the bracket. From there the competition becomes single elimination, with 63 matches in total and six rounds to survive, the first of which accounts for 32 of those matches on its own. Individual matches are played over the best of five sets of three arrows, two set points to whoever wins a set and six set points to close the match, which is the arithmetic I have written about elsewhere.

This article concerns only the first link in that chain: what the ranking measures, and how precisely. Everything downstream, from the set format to the shoot-off arrow, comes later and is not needed to answer the question.

What "the best" means here

The criterion has to be stated before the numbers, because without it the percentages mean nothing. In the work these figures come from, the best archer is defined in one way and a fairly narrow one: the archer who, stepping onto the line, expects the highest average score. No other quality enters the definition, and it is left out by choice rather than by oversight.

So nerve in a close finish, the ability to read a turning wind, coolness on a shoot-off arrow, and the knack of recovering after a lost set all sit outside the calculation, and those are qualities archers recognise perfectly well and nobody considers irrelevant. The consequence is precise: anyone who rejects that definition of ability should not be correcting the numbers that follow, they should be rejecting the question the numbers answer. That is a legitimate position, as long as it is clear that it is the position being taken.

The three numbers

The way to measure what a ranking is worth is to have archers of known ability shoot, meaning archers whose expected score is fixed in advance, and then count how often the order produced by their 72 arrows matches the true order. Repeat that over sixty thousand qualification rounds for each scenario and the answer stops depending on the luck of any single day. For a compact field, with 64 archers spread between 690 and 665 points, the results are these.

What the ranking is asked to do How often it manages
Put the best archer in the field first14.2 times in 100, a little over once in seven
Put the best eight in the first eight places4 times in 10,000
Put each archer in the place they deservemean error of 8.6 places

The first row has a translation that carries the order of magnitude better than the percentage does. On a circuit of seven competitions a year, the number one seed is the strongest archer present exactly once, and in the other six somebody else is, without anyone noticing, because the bracket does not say so.

The second row deserves a slow read, because it is the one that surprises. That the eight archers sitting in the first eight slots really are the eight best in the field happens four times in ten thousand, which is fewer than one competition in two thousand. Every time protecting the seeds comes up for discussion, then, what is being protected is eight archers who in practice are almost never the right eight.

The third row says the same thing in a form that is easier to use when you are looking at a real ranking. A mean error of eight and a half places, on a grid of 64, means that an archer entering the bracket as number 50 may perfectly well be worth twentieth, and may meet somebody in the first round they should never have met. It works the other way round too, which is the half worth remembering when you find your own name near the bottom of the list.

Why a good measurement makes a bad ranking

The explanation is not a defect in the qualification round, because as a measuring instrument 72 arrows work very well indeed. Set to compare two archers 10 points apart, they order them correctly 95 times in 100, which is an error once in 20, and for a sporting measurement that is an excellent result that stands comparison with any other selection criterion in use.

The trouble comes from the question put to the instrument. In a compact top-level field all 64 competitors sit inside 25 points, which means two neighbours in the ranking are separated on average by four tenths of a point rather than by ten. Over a gap that size the same 72 arrows point to the better archer 52.9 times in 100, which is barely distinguishable from a coin. And since a ranking of 64 names rests on 63 comparisons between neighbours, that uncertainty accumulates down the whole list instead of staying confined to one pair.

The image the essay uses is a kitchen scale. The scale works, and it will tell one bag of flour from another that weighs 100 g more without hesitating. What it is being asked to do here is put 64 bags in order when they differ from one another by a gram, and no kitchen scale can do that however good it is.

Key point

The instrument works. The question is finer than the instrument.

Seventy-two arrows order two archers 10 points apart correctly 95 times in 100. Between two neighbours in a compact ranking, four tenths of a point apart, the same measurement gets it right 52.9 times in 100. What fails is asking a qualification round to sort 64 archers packed inside 25 points.

Four tenths of a point, seen on the target

To see how out of proportion that request is, it helps to turn the points into something physical, namely the width of the group an archer leaves on the face. An archer worth an average of 700 points shoots inside a group 3.78 cm wide. At 690 the group is 4.46 cm, at 680 it is 5.14, at 670 it is 5.81, and at 660 it is 6.49. The lower the score, the wider the group, and the relationship between the two scales is almost perfectly regular.

FIG · 01 Ten points are seven millimetres of group Group size matching each mean score over 72 arrows at 70 metres 4 cm 6 cm 8 cm 8.527.506.49 5.815.144.463.78 680 to 690: 0.68 cm 630645660 670680690700 MEAN SCORE OVER 72 ARROWS Calculated from the geometry of the 122 cm face, not measured on the field. The figure is the size of the whole group.
Fig. 01 The scale of points and the scale of centimetres run in step. Going from 680 to 690 points means tightening the group by almost seven millimetres, while the four tenths of a point that separate two neighbours in the ranking are worth a little over a quarter of a millimetre.

With that translation to hand, 10 points out of 720 stop looking like a detail. Going from 680 to 690 means tightening the group by almost 7 mm, which is about 15% of its width, and starting from 700 the same ten points are worth 18%. Two archers 10 points apart differ by something a decent instrument recognises nearly every time, and 72 arrows do recognise it, in 95% of cases.

Four tenths of a point, by contrast, are worth a little over half a percentage point of group size, which on a group 4.5 cm across comes to a little over a quarter of a millimetre. Asking 72 arrows to separate two groups that differ by a quarter of a millimetre is asking for a resolution that number of shots does not have, and the ranking answers with the only thing it can produce, which is an order decided largely by chance among neighbours.

The toll of seeding

An error of eight and a half places does not stay confined to the first round, because the bracket goes on trusting the ranking for the whole competition. The essay measures that cost separately and calls it the toll of seeding: you run the same structure twice, once with the true seedings and once with the ones the qualification round produced, and look at what is lost between the first run and the second.

The result is that every structure pays in proportion to how often it consults the ranking. The standard bracket, which asks once at the start and then never looks again, loses seven tenths of a percentage point on the probability that the title ends up with the best archer. Quarterfinal repechage, which consults it again to place the archers who have been knocked out, loses one and six. Reseeding, which rebuilds the pairings every round and therefore asks the seeding six times, loses four and eight.

There is an irony inside that result worth keeping in mind whenever formats come up for discussion, because the structures that look on paper like the best correctors of bad pairings are also the ones that lean hardest on an imprecise measurement, and therefore pay most for doing so. That does not make them worse in absolute terms, since the advantage they bring can outweigh the toll. It does explain why the net balance always comes out thinner than the logic of the format would lead you to hope.

Why a fairer ranking is not enough

From all this it would be natural to conclude that the fix is a better qualification round. The essay tests that route as well, and the result is the most instructive part of the whole volume: raise the agreement between the ranking and the true order from 0.81 to 0.99, which is to say buy an almost perfect measurement, and the standard bracket goes from 9.7% to 10.4% of titles awarded to the best archer. Inside the error of the measurement, that means gaining nothing.

The reason is that the bracket does not consult the ranking to find out who is best. It consults the ranking to decide who meets whom. What actually matters is how it distributes opponents across all 63 matches, and on that distribution a perfect ranking changes little, because it still sends half the field to eliminate each other in the first round. The same data, applied to a structure with reseeding, take it from 11.1% to 14.0% instead: the same improved measurement returns four times as much the moment the structure is able to use it.

The order of operations, then, is the reverse of the one most often proposed. Change the structure first, improve the measurement second. Doing it the other way round, which is also the easier decision because it leaves the rules of the finals untouched, produces no measurable result.

What changing the criteria buys

Inside the qualification round there are two smaller levers, and they have been measured too because they come up often on shooting fields. The first is the set of criteria used to break ties on score, meaning the count of Xs and then the count of tens. Against a ranking in which ties are drawn by lot, adding Xs takes the probability that the top qualifier really is the best from 14.0% to 14.2%, and adding the count of tens on top of that moves it not at all.

The second lever is more radical and consists of ordering the qualification round by the mean distance of the arrows from the centre rather than by score, recovering the information the scoring rings throw away every time they round an impact to the value of its zone. That one works: the top qualifier becomes the best archer 15.9 times in 100 instead of 14.2, and the mean error falls from 8.6 places to 7.7. On the title, though, the effect stays at half a percentage point, for the reason given above, which is that a better measurement is of little use to a structure that cannot do anything with it.

It depends who is in the competition

All the figures so far are for a compact field, and the same qualification round performs a great deal better when the field is put together differently. This is the correction that keeps the argument honest, because without it the whole thing would read as an accusation against the measurement rather than a description of its limit.

Shape of the field Top qualifier is the best Top eight are right Mean error
Compact, 690 down to 665 points14.2%0.04%8.6 places
Bimodal, 16 archers from 692 to 680 plus 48 from 655 to 63024.8%1.9%6.8 places
Spread, 695 down to 625 points32.4%5.1%4.1 places

In a spread field, where 70 points separate first from last instead of 25, the mean error more than halves and the top qualifier is the best archer almost once in three. The practical consequence is for anyone reading the results of a regional competition rather than a World Cup final: the wider the field, the more information a finishing position carries. The essay sets the opposite observation alongside it, which is that in a competition like that the best archer is visible to the naked eye, so the measurement matters least exactly where it works best.

The row that carries weight is still the first, because a top-level field is compact by definition. You get there by selection, and selection is precisely the process that removes the distant archers and leaves the ones who resemble each other. The higher the level, then, the less informative the qualification ranking becomes rather than the more, which is the opposite of how it is read.

Doubling the arrows does not fix it

The obvious proposal is to shoot more, and it moves in the right direction without changing the substance. Take the qualification round to 144 arrows instead of 72 and, in a compact field, the top qualifier really is the best archer in 18.9% of cases instead of 14.2%, the top eight come out right 0.3% of the time, and the mean positional error falls from 8.6 places to 6.5.

That is a real gain, and it costs half a competition day more at every event on the calendar. What remains is still a ranking that puts the best archer first fewer than one time in five, and the reason it returns so little is arithmetic: the precision of a repeated measurement grows as the square root of the number of trials, so halving the uncertainty would take four times the arrows, and bringing the qualification round up to the reliability everyone already credits it with would take a number of shots no calendar could hold.

Who actually wins the title

Moving from the ranking to the whole tournament changes the question: how often does the gold end up with the best archer in the competition? In the model, with the structure currently in force, that happens 9.7 times in 100, which is one competition in ten.

The two figures have to be kept carefully apart, because they measure different things and are very easy to confuse. The 14.2% is about the ranking, meaning how often the top qualifier is the best archer in the field. The 9.7% is about the whole tournament, meaning how often the title ends up in the hands of the best archer after six rounds of single elimination. The first is a property of the measurement, the second a property of the competition.

The lever that actually moves the second figure is not the qualification round but where the seeding comes from. Use the season rankings in place of the score from the day, and rebuild the pairings at every round, and the title goes to the best archer 13.7 times in 100 instead of 9.7: four percentage points, which in this field is a large improvement and costs a deep change to the rulebook. Even then, in six competitions out of seven the title still goes to somebody who is not the best archer present, and that is a property of single elimination rather than a fault anyone can correct.

What the calculation does not look at

These numbers come out of a modelling study and not from a collection of real results, and the essay declares it on the first page: no competition data was used. The archers are constructed profiles defined by their expected score, and the three fields of competitors are assembled deliberately using two declared dials, the mean level and how far apart the competitors are. Even the 25 points separating first from last in a compact field are a stated assumption rather than a measurement taken on World Cup fields.

On the simulated values the essay recommends reading two significant figures rather than three, and treating differences below half a percentage point as absent. When I write 9.7 or 13.7, those figures carry a margin of a few tenths, and comparisons have to be made with that margin in mind. The same quantity, measured by different runs of the same model, appears elsewhere with slightly different digits, in the way that weighing yourself on two consecutive mornings gives two numbers.

There is also something I am asked about often that the work does not say: how often the number one seed actually wins the competition, or goes out in the first round. That figure is not there, I have not derived it by subtraction from the others, and you will not find it in this article, because the two percentages I have quoted are about who is the best archer in the field and not about who wins. It is worth saying plainly, since it is exactly the question a headline like this one invites.

Finally, the direction of the limitations is declared and it all runs one way: four of the model's five simplifications push towards a smaller ability to tell archers apart, not a larger one. Wind, for instance, widens everyone's groups and pushes the positional error above 8.6 places. That the real field sits below the calculated values remains an extrapolation rather than a measurement, but none of the known simplifications works in the direction of making the ranking more reliable.

What to do with a ranking

For the archer the practical consequence is that your own finishing position is a noisy label, and so is your opponent's. If you came in twelfth rather than eighth, those four places are worth a little over a point and a half of expected score in a compact field, which sits inside the noise of the measurement and establishes no stable difference in precision between you and the archers ahead of you.

The reverse holds too, and in that form it is the useful half to remember the evening before a match: meeting the number one seed is not the same as meeting the best archer in the field, because that number is right a little over once in seven. First against sixty-fourth, in a compact field, is not the gulf the bracket suggests, and going into that match convinced otherwise is a way of giving away half a set before it starts.

For anyone coaching, the translation is that a finishing position from a single competition is not material to build an assessment on, while the series of finishing positions across a season is. The reason is the same one that makes an average over several competitions tighten the uncertainty around a score, and I have written about it in how many points it takes to say one archer is better than another. The same 72 arrows that say little on their own say a great deal once they become 300.

Where this goes next

The volume these calculations come from goes well beyond the qualification round. It measures alternative bracket structures, repechage, double elimination, match formats, and season rankings, and it arrives at a price list of what each choice of rule costs and what it returns.

The result that holds it all together is the mismatch between what gets discussed and what actually weighs. The match format, which is what people argue about most, moves outcomes by a few percentage points. The instrument that decides who meets whom leaves the order of almost every neighbouring pair undetermined, and nobody questions it, because it produces a list of numbers that looks like a verdict.

Questions I get asked most

How does an Olympic archery competition work?
It opens with a qualification round of 72 arrows at 70 metres on a 122 cm face, for a maximum of 720 points. The totals produce a ranking whose only job is to set the pairings, with first against sixty-fourth and second against sixty-third. The competition then moves to single elimination: 64 archers, 63 matches, and six rounds that all have to be won to take the title.

How reliable is the qualification ranking?
Not very, in a top-level field. The top qualifier really is the best archer in the field 14.2 times in 100, the top eight really are the best eight four times in ten thousand, and the mean positional error is 8.6 places on a grid of 64. In a more spread-out field the results improve considerably: the top qualifier is the best archer in 32.4% of cases and the error falls to 4.1 places.

Why are 72 arrows not enough to make a fair ranking?
Because the question is finer than the instrument. Two archers 10 points apart are ordered correctly 95 times in 100, but in a compact field two neighbours in the ranking are four tenths of a point apart, and over that gap the same arrows get it right 52.9 times in 100. Sixty-four archers packed inside 25 points are too close together to be sorted by a measurement of that resolution.

Does the best archer always win the competition?
No, and not even close. In the model the title ends up with the best archer in the field 9.7 times in 100. Improving the ranking without touching the bracket would take that to 10.4%, while seeding from the season rankings and rebuilding the pairings every round would take it to 13.7%. Even at best, in six competitions out of seven the title goes to somebody who is not the strongest archer present.

Where these numbers come from. Every figure in this article comes from the essay When Nobody Won: ties, competitions and seasons in archery, published in full in the Lab, where the calculation is worked through step by step, and every assumption of the model is declared. The practical readings for archers and coaches are mine: the essay deals only with the calculation. This is a modelling study: no competition data was used, and the figures hold under the assumptions the essay declares.

Go to the bibliography
The seeding says little. The group says everything.

Ten points of score are seven millimetres of group.

What a ranking cannot measure is perfectly visible on the target and in the movement that produced it. Send me a video of your shot and I will send back the biomechanical reading with the data and the priorities to work on.