What a pedigree can't tell you about a foal
Every inbreeding figure a breeder works from is a prediction rather than a measurement — an average over all the foals a mating could have produced, not a statement about the one standing in the field.
Take two full brothers. Same sire, same dam, born two years apart.
Any pedigree software in the world will give them the same inbreeding coefficient. It has to: they have identical ancestry, and the coefficient is calculated from ancestry. Yet the two horses do not carry the same genome, and they are not inbred to the same degree. One may have inherited noticeably more duplicated ancestral material than the other, and nothing in the paperwork will ever tell you which.
This is not a flaw in anybody's software. It is what the number has always been, and it is worth understanding properly — particularly if you breed in a population small enough that these decisions matter.
What the coefficient actually is
Sewall Wright gave us the method for calculating inbreeding from a pedigree in 1922. It answers a specific question: given this ancestry, what proportion of an animal's genome would you expect to be identical by descent?
The word doing the work is expect. It is an average across all the foals that mating could theoretically produce. Inheritance itself is a lottery — chromosomes recombine, and each foal draws a different hand from the same deck. Wright's figure describes the deck. It does not describe the hand.
Three things follow, and each of them bites harder as a population gets smaller.
Founders are assumed unrelated. Every pedigree analysis starts somewhere, and whatever animals sit at the base of the studbook are treated as unrelated to one another and not inbred themselves. In a breed founded from a genuinely broad base, that is a tolerable approximation. In a breed founded from a handful of local animals who were almost certainly cousins, it is a fiction — and it is a fiction that always errs in the same direction, downwards. The team behind the recent Thoroughbred genome work put it more bluntly than I would dare: the assumption that all founders of a population are unrelated is, they write, unarguably violated.
Depth decides the answer. A coefficient calculated over five generations and one calculated over twenty-five are different numbers describing the same horse. Everything before the cut-off is invisible and counted as zero.
Recombination adds noise. This is the full-brothers problem, and it has been quantified. Simulation work published in Heredity in 2017 found that for offspring of a full-sibling mating — pedigree coefficient 0.25 — the actual proportion of genome inherited identical by descent varied with a standard deviation of about 0.045 when modelled on a human genetic map. If that spread is roughly normal, about one animal in three sits more than 0.045 away from the number on the page. The same work found the relative error grows as the inbreeding becomes more distant: on that same human map, by the time you reach second cousins the standard deviation is around three-quarters of the coefficient itself, and on a zebra finch map it exceeds it.
Those two maps are the ones the authors modelled. Nobody has run the equivalent for the horse, which has thirty-one autosomes against the human twenty-two — and since the finch result came out at nearly twice the human one, genome architecture clearly moves the answer. Treat the horse figure as unknown rather than assumed.
What genomics measures instead
The alternative is to stop predicting and start looking.
When a stretch of chromosome is inherited from a common ancestor down both sides of a pedigree, it arrives as a long, unbroken run in which the animal carries the same version twice over. These are runs of homozygosity, and modern genotyping can find them directly. Add them up, divide by genome length, and you have FROH — not an expectation, but a measurement of what this particular animal actually inherited.
There is a bonus. Recombination chops these runs shorter with every generation, so length carries a date stamp: long runs mean recent common ancestors, short runs mean old ones. A pedigree coefficient is one number. A genomic profile tells you when.
What happens when you measure both
The honest summary is that the two agree on direction and disagree on magnitude, and the disagreement is not small.
The clearest illustration comes from 185 North American Thoroughbreds whose whole genomes were sequenced and published in Scientific Reports in 2024. Measured genomically, their inbreeding averaged 0.266 and 0.283 in the two sampled groups. Set directly against five-generation pedigrees for the same horses, the genomic figure was 0.240 and the pedigree figure 0.010. The two correlated at r = 0.40.
Two caveats before anyone reaches for that as a scandal. These horses were not a random sample — one group was deliberately chosen to maximise pedigree diversity, which will have pulled the pedigree figure down. And five generations is shallow. Earlier work tracing Thoroughbred pedigrees back twenty to thirty generations put inbreeding at around 0.139, and a study of Australian Thoroughbreds averaging nearly twenty-three generations per horse reached 0.143. So the pedigree answer for the same breed moves from 0.010 to about 0.14 depending only on how far back you look — against a genomic figure of roughly 0.27, which captures homozygosity older than the studbook itself.
That is the real lesson. The pedigree number is not wrong so much as it is relative to a starting line, and anyone quoting one should know where their starting line is.
The same paper contains a finding that ought to interest anyone planning a mating. Comparing horses pairwise, the proportion of the genome two animals actually shared as homozygous correlated only r = 0.33 with their pedigree relationship. Two horses can each be substantially inbred without being inbred in the same places, which is what determines whether a recessive meets its match. As the authors put it, it is difficult to predict the true proportion of homozygosity in the genome of a foal from pedigree relationship alone.
Across other breeds the pattern repeats:
- Franches-Montagnes. Pedigree 6.84 per cent against genomic 12.15 — near enough double, in a breed with more than twenty generations of pedigree behind it. (500 kb minimum run; Pearson r = 0.56.)
- Icelandic horses. Pedigree 0.03 against genomic 0.20, in a breed with genuinely good records — over eight generation-equivalents deep. (100 kb minimum run; Pearson r = 0.57.)
- Slovenian Lipizzans. Pedigree 0.037 against genomic 0.122. (1 Mb minimum run; Spearman 0.56.)
- Polish cold-blooded horses in a national conservation programme. Rank correlation with genomic measures was 0.44 overall, rising to 0.67 among animals with fifteen recorded generations — though that subgroup figure rests on a thin sample and should be held loosely. The same literature notes imported stallions known to be related appearing in the database as unrelated, and therefore assigned an inbreeding of zero.
Read the brackets, because they matter more than they look. Those genomic figures are not comparable with one another: each study set its own minimum run length, and that setting moves the answer substantially. Nor are the correlations all the same statistic — some are Pearson, some Spearman.
What survives the caveats is the direction. In almost every case the genome reports more inbreeding than the paperwork does. What does not survive is the comfortable assumption that deeper records fix it: the Icelandic studbook is deep and the Franches-Montagnes deeper still, and both still show roughly a doubling. Better records improve the correlation between the two measures. They do not reliably close the gap.
Where I should be careful
It would be easy to write the rest of this as pedigrees bad, genomics good, and that would be an overcorrection.
Genomic figures come with their own settings. Change the minimum run length you count and the answer moves a long way — in the Icelandic and Exmoor study, going from a 100 kb to a 500 kb threshold took the figures from 0.20 to 0.08 and from 0.27 to 0.12 respectively. On a like-for-like basis, some of the rankings above would reverse. The Thoroughbred work also found that when SNP arrays were simulated from full sequence data, the inbreeding attributed to the longest runs came out four to five times higher than sequencing showed — though the same comparison found array and sequence estimates correlated above 0.98, so the problem is absolute calibration rather than rank order.
FROH is a measurement, but it is a measurement with dials.
More importantly: pedigree management demonstrably works. It has saved breeds. Nobody should read any of this as a reason to stop doing it.
The study one breed has never had
Which brings me to the case I keep returning to.
The Cleveland Bay is among the most carefully documented rare horse breeds anywhere. Its first studbook was published in 1885 and carries retrospective pedigrees reaching back to 1732. It has been through a proper population-genetic analysis, and the results were sobering: work published in PLOS ONE in 2020 found that 91 per cent of stallion lines and 48 per cent of dam lines have been lost, that just three ancestors account for half the genome of the living population, that 70 per cent of the present female population descends from three founder females, and that every paternal line traces back to a single founder stallion. At the time of that study the Rare Breeds Survival Trust listed the breed as critical — one of five equines on its Watchlist with fewer than 300 registered adult breeding females.
And the response to that analysis worked. Managed matings have been the breed's conservation strategy for two decades, with measurable results — the story of how that happened is one of the better rare-breed rescues of this century.
Here is the gap. Everything published on Cleveland Bay genetics rests on pedigree records and microsatellites: fifteen markers in the 2020 study, fifteen again in an independent American survey published in Diversity in 2019, plus a separate analysis of mitochondrial variation. Fifteen markers is a tool for verifying parentage and comparing breeds. It is not a genome scan, and it cannot find runs of homozygosity.
As far as I can establish, no SNP-array or runs-of-homozygosity study of the Cleveland Bay has been published. The breed is absent from the panels behind the EquineSNP50 and the high-density Axiom array, from every published multi-breed equine homozygosity survey I can find, and from the recent survey of native British and Irish ponies — which sampled eleven pony populations and, being ponies, did not include it. Its only appearance in a sequencing study that I can trace is a single male in a Y-chromosome capture panel, which yields no autosomal information and cannot speak to this. I would be glad to be corrected by anyone who knows of such a study.
So the honest position is this. We know a great deal about how related Cleveland Bays are on paper, and effectively nothing about how much of that relatedness actually landed in the animals now alive. Given how consistently the two diverge in every other breed where both have been measured — and given that this is a small, historically bottlenecked, closed population, the exact profile where divergence is largest — that is a question worth answering rather than assuming.
I am not suggesting the answer would overturn the current strategy. The likelier outcome by far is that it confirms it, adds resolution, and shows which individual animals carry more duplicated genome than their paperwork implies. In the breeds where this work has been done, that is roughly what it has produced: not a reversal, but a sharper instrument.
It is also not a job for breeders, and I want to be clear about that. A genomic survey of a breed is a university-and-society undertaking — sampling, funding, ethics, analysis, publication. It is exactly the kind of collaboration that produced the FIS carrier test for the Fell and Dales, which I wrote about earlier today and which came out of the Animal Health Trust at Newmarket working with the University of Liverpool, rather than out of anybody's yard.
Why we are writing about this
Because we sell access to breeding animals, and I would rather be honest about the limits of the information we display than quietly imply it is more precise than it is.
When a listing carries a pedigree-based figure, that figure is a well-founded estimate of what a mating is likely to produce on average. It is not a measurement of the animal in front of you, and in a small closed population it is probably an underestimate. We will keep publishing those figures, because they are the best available and because the alternative is publishing nothing. But we should say what they are.
The breeds we care most about are the ones where the paper number and the real number are furthest apart. That is not an argument for trusting the paperwork less. It is an argument for someone, eventually, going and looking.
Further reading
- Bailey, E., Finno, C.J., Cullen, J.N., Kalbfleisch, T. & Petersen, J.L., "Analyses of whole-genome sequences from 185 North American Thoroughbred horses, spanning 5 generations", Scientific Reports 14:22930, 2024 (open access)
- Knief, U., Kempenaers, B. & Forstmeier, W., "Meiotic recombination shapes precision of pedigree- and marker-based estimates of inbreeding", Heredity 118:239–248, 2017
- Kardos, M., Luikart, G. & Allendorf, F.W., "Measuring individual inbreeding in the age of genomics: marker-based measures are better than pedigrees", Heredity 115:63–72, 2015
- Dell, A., Curry, M., Yarnell, K., Starbuck, G. & Wilson, P.B., "Genetic analysis of the endangered Cleveland Bay horse: A century of breeding characterised by pedigree and microsatellite data", PLOS ONE 15(10):e0240410, 2020 (open access)
- Dell, A. et al., "Mitochondrial D-loop sequence variation and maternal lineage in the endangered Cleveland Bay horse", PLOS ONE 15(12):e0243247, 2020 (open access)
- Dell, A. et al., "16 years of breed management brings substantial improvement in population genetics of the endangered Cleveland Bay Horse", Ecology and Evolution 11(21):14555–14572, 2021 (open access)
- Khanshour, A.M., Hempsey, E.K., Juras, R. & Cothran, E.G., "Genetic Characterization of Cleveland Bay Horse Breed", Diversity 11(10):174, 2019 (open access)
- Winton, C.L. et al., "Genetic diversity within and between British and Irish breeds: The maternal and paternal history of native ponies", Ecology and Evolution 10(3):1352–1367, 2020 (open access)
- Sigurðardóttir, H. et al., "Genetic diversity and signatures of selection in Icelandic horses and Exmoor ponies", BMC Genomics 25:772, 2024 (open access)
- Luštrek, B. et al., "Comparing Genomic and Pedigree Inbreeding Coefficients in the Slovenian Lipizzan Horse as a Case Study for Small Closed Populations", Animals 15(19):2774, 2025 (open access)
- Polak, G. et al., "Suitability of Pedigree Information and Genomic Methods for Analyzing Inbreeding of Polish Cold-Blooded Horses Covered by Conservation Programs", Genes 12(3):429, 2021 (open access)
— Rene
Pieces along the same line
The first item on the Sealyham's health list is not a test
The Royal Kennel Club's Health Standard names two things as good practice for the Sealyham Terrier — the minimum, not the aspiration.
Read the piece →The Clumber Spaniel stud list that publishes its carriers
Ten years of Kennel Club registrations for the Clumber Spaniel — 171, 265, 280, 175, 188, 285, 232, 223, 155, 160 — with the two lowest years at the end of the run.
Read the piece →Want more like this in your inbox?
The GenoVaq journal publishes long-form pieces for breeders and buyers — welfare, health-testing, breeding decisions, marketplace mechanics. New writing every week or two.