Note: The ethical judgments on this page refer exclusively to the action — never to the person who performs it or who came into existence through it. Cf. Note on Ethical Judgments.
Algorithmic embryo assessment is a ranking attribution about a human embryo generated by an artificial system. It rests on image features — still images or time-lapse recordings from the incubator — and on a statistical model that has learned from large image holdings which appearance went along with which outcome. The result is a number, and from it an order.
Ontologically this attribution as such is not yet a choice. It becomes an action upon persons only where transfer and discarding follow from it — that is algorithmic embryo selection. The separation is no hair-splitting: on it hangs which objections strike whom.
Ontological classification
- is a subclass of: algorithmic assessment — the attribution about a person generated by an artificial system
- is thereby: a temporal attribution and expressly not essence-determining
- is presupposed by: algorithmic embryo selection
- is typically generated in: automated reproduction laboratory
- presupposes: in vitro fertilization and hence several simultaneously available embryos
- is objected to by: the objection from algorithmic bias and the objection from lacking explainability — both consequence-based objections
The decisive determination the concept inherits from its superclass: what an algorithm attributes to a person is a statement about a point in time and about a constellation of characteristics, not about what this person is. This holds for the credit rating of an adult just as for the score about a blastocyst. The assessment becomes person-relevant not through the possibility of its being wrong, but through the discrepancy between the status of the attribution and the reach of the decision based upon it: a temporally limited, non-essence-determining assessment becomes the ground of a decision about existence.
What the score actually predicts
The term “assessment” promises more than the procedures deliver. The widespread systems rank by the probability of implantation, or of a fetal heartbeat — not by live birth, not by chromosomal status, not by the health of the child. Whoever reads the score as a statement about which embryo is “the best” reads into it something that was not measured.
Models built expressly for the prediction of chromosomal abnormalities train against the findings of preimplantation genetic diagnosis (PGT-A). In principle they can therefore be no better than this training target, which is itself error-laden — mosaic findings and misclassifications included. To this is added a problem of circularity: chromosomal abnormalities accumulate with maternal age, and this leaves its traces in the morphology. An image model can thus achieve high hit rates by in fact representing the woman’s age, without ever having derived any chromosomal information from the recording. The number is then correct and the ground is other than the one asserted.
The reference standard is itself imprecise
The comparison is made against morphological assessment by embryologists according to the Gardner scheme — degree of expansion, inner cell mass, trophectoderm, combined into fifty-four categories. Agreement between different assessors lies there at a kappa value of around 0.63 to 0.66. This is the actual starting point of the manufacturers’ standardization argument: not that the machine is particularly good, but that the human being is not particularly reproducible here.
The argument is to be taken seriously, but it carries only so far as reproducibility is a value in itself. An assessment can be perfectly stable and at the same time perfectly uninformative.
What the evidence shows
The decisive trial comes from Illingworth, Venetis, Gardner, Nelson and others (Nature Medicine 2024): fourteen clinics, 1,066 patients, double-blind, an established deep learning system against conventional morphological assessment. Result: clinical pregnancy 46.5 versus 48.2 percent, live birth 39.8 versus 43.5 percent. The prespecified noninferiority was not reached; the authors state verbatim that noninferiority for the clinical pregnancy rate could not be shown. The only demonstrated advantage was a saving of time. The study was manufacturer-funded — the result came out against the sponsor’s interest, which increases its evidential force.
The TILT trial (The Lancet 2024) likewise found no higher live birth rates for time-lapse procedures. The British supervisory authority HFEA lists time-lapse incubation with automated analysis on the worst step of its traffic-light scale and justifies this not by a safety risk but by the absence of an effect on the treatment outcome.
More instructive than any debate about efficacy are two studies on internal consistency. Zaninovic and others (F&S Science 2024) had a hundred cycles ranked by five embryologists and eight commercial systems: agreement between embryologists at Kendall’s τ = 0.70, between human and machine at 0.53, between the machines among themselves only at 0.47; two of the eight systems lay at chance level. Thirumalaraju and others (Fertility and Sterility 2026) trained the same model repeatedly on the same data and obtained concordances around W = 0.35, critical error rates of 12.4 to 17.3 percent and ROC areas between 49.5 and 71.8 percent — the worst replicate lies below chance.
A further finding concerns a gap: to date there exists no published study that breaks down the performance of such a system by ethnicity. The charge of bias is thereby neither substantiated nor refuted — it is untested.
Ethical assessment
The assessment here follows not an authority but the line of justification of the ontology of personhood. Ontological dignity is the inalienable, objective worth of the person and the sufficient ground of the personalistic norm: the person is to be affirmed for her own sake. This norm is violated, among other ways, by forgetfulness of the person — the affirmation of a person not for her own sake but on account of her characteristics.
The assessment taken by itself is no such violation. A number about a recording is a description; it kills no one and chooses no one. Nor does the circumstance that an artificial system generates it change anything: the system carries out what is assigned to it and is in that respect value-neutral. The break lies at the point where the attribution becomes the ground. Then an embryo — a person in the First Dimension — is affirmed on account of a characteristic and rejected on account of a characteristic. That is the matter of eugenic selection, and it obtains independently of how reliably the procedure works.
What algorithmic assessment adds to this matter is not a new moral quality but a displacement of attributability: the responsibility gap. In regulated operation it appears regularly as nominal oversight — the responsible party is named and signs off, but for want of explainability cannot examine what he answers for.
The objections thereby fall cleanly into two classes. Biased training data, lacking explainability, inadequate information of the patient, diffuse attribution — these are consequence-based objections and curable by a procedure. The objection to the discarding by characteristics is an act-based objection and curable by no procedure. A governance framework answers without exception the first class. Where it is offered as an answer to the second, the question is not answered but passed over.
What stands objected to before any assessment
Algorithmic assessment constructively presupposes in vitro fertilization, since it needs several simultaneously available embryos in order to form a ranking at all. In the ontology of personhood this is carried, by way of artificial fertilization, as an intrinsically evil act.
The evaluation therefore reaches the attribution along the path of presupposition and not by inheritance — and it applies even where the algorithm works reliably, explainably and under supervision. According to Donum vitae II, B, 5, even homologous fertilization without any destruction of embryos remains illicit. Whoever begins only at the quality of the attribution is examining a procedure whose admissibility would have to be settled beforehand.
The strongest objection to this assessment
One may reply: the assessment does not choose at all, it only orders. In in vitro fertilization several embryos regularly come into being, and for medical reasons as a rule only one is transferred. Ranking therefore must take place; the only question is whether a human being or a machine ranks. Whoever censures algorithmic assessment in truth censures the antecedent practice — and does so in the wrong place, for the algorithm makes the situation of the supernumerary embryos no worse by anything. It is at best more accurate and at worst just as good as the embryologist.
The objection strikes, and it strikes harder than it looks at first: the critique conducted here in fact cannot be fastened on the algorithm, but only on the ranking itself. The answer must therefore be twofold. First: yes — the act-based objection is directed against the discarding by characteristics, not against its technique. Precisely that is the reason why automation is treated here as a withdrawal of the personal performance and not as a new wrong. But second: the algorithm is not neutral toward the practice it serves. It lowers its costs, raises its throughput, lends it the appearance of objectivity and makes the ranking repeatable, recordable and scalable. A practice whose justification hangs on its unavoidability is not strengthened in its justification by a tool that makes it effortless, but weakened.
And in between the finding of fact stands firm: what is here made the ground of a decision about existence is a prediction of the probability of implantation, measured against an imprecise standard, at odds with its own kind, and without demonstrated superiority in the large trial. The attribution is temporal. The decision is not.
See also
- Algorithmic Embryo Selection — the action that builds on the assessment
- Automated Reproduction Laboratory
- Nominal Oversight, Responsibility Gap
- Act-Based Objection, Consequence-Based Objection, Governance Framework
- Eugenic Selection, Preimplantation Genetic Diagnosis (PGD)
- In Vitro Fertilization, Embryo Transfer, Supernumerary Embryo
- AI-Algorithmic Arrangement, Forgetfulness of the Person
- Embryo, Ontological Dignity
Sources: Generated by querying the ontology of personhood. Research as of 29 July 2026.
Further sources:
- Illingworth, P. J., Venetis, C., Gardner, D. K., Nelson, S. M. et al. (2024): Deep learning versus manual morphology-based embryo selection in IVF: a randomized, double-blind noninferiority trial. Nature Medicine 30(11): 3114–3120.
- Zaninovic, N., Sierra, J. T., Malmsten, J. E., Rosenwaks, Z. (2024): Embryo ranking agreement between embryologists and artificial intelligence algorithms. F&S Science 5(1): 50–57.
- Thirumalaraju, P., Kanakasabapathy, M. K. et al. (2026): Stability and reliability of artificial intelligence models in embryo selection for in vitro fertilization. Fertility and Sterility 125(2): 277–286.
- Bhide, P. et al. (2024): Clinical effectiveness and safety of time-lapse imaging systems for embryo incubation and selection in in-vitro fertilisation treatment (TILT): a multicentre, three-parallel-group, double-blind, randomised controlled trial. The Lancet 404(10449): 256–265.
- Human Fertilisation and Embryology Authority: Rating of IVF add-ons (traffic-light scale), entry “Time-lapse imaging with automated analysis”.
- Gardner, D. K. & Schoolcraft, W. B. (1999): In vitro culture of human blastocysts. In: Jansen, R. & Mortimer, D. (eds.): Towards Reproductive Certainty. Carnforth: Parthenon Press, pp. 378–388. — Grading scheme by degree of expansion, inner cell mass and trophectoderm.