And even if it is right, an AI agent can’t complement the information it provides with the knowledge physicians gain through experience, says fertility doctor Jaime Knopman. When patients at her clinic in midtown Manhattan bring her information from AI chatbots, it isn’t necessarily incorrect, but what the LLM suggests may not be the best approach for a patient’s specific case.
For instance, when considering IVF, couples will receive grades for viability for their embryos. But asking ChatGPT to provide recommendations on next steps based on those scores alone doesn’t take into consideration other important factors, Knopman says. “It’s not just about the grade: There’s other things that go into it”—such as when the embryo was biopsied, the state of the patient’s uterine lining, and whether they have had success in the past with fertility. In addition to her years of training and medical education, Knopman says she has “taken care of thousands and thousands of women.” This, she says, gives her real-world insights on what next steps to pursue that an LLM lacks.
Other patients will come in certain of how they want an embryo transfer done, based on a response they received from AI, Knopman says. However, while the method they’ve been suggested may be common, other courses of action may be more appropriate for the specific patient’s circumstances, she says. “There’s the science, which we study, and we learn how to do, but then there’s the art of why one treatment modality or protocol is better for a patient than another,” she says.
Some of the companies behind these AI chatbots have been building tools to address concerns about the medical information dispensed. OpenAI, the parent company of ChatGPT, announced on May 12 it was launching HealthBench, a system designed to measure AI’s capabilities in responding to health questions. OpenAI says the program was built with the help of more than 260 physicians in 60 countries, and includes 5,000 simulated health conversations between users and AI models, with a scoring guide designed by doctors to evaluate the responses. The company says that it found that with earlier versions of its AI models, doctors...








English (US)