Breaking news

AI Model Matches And At Times Exceeds Doctors In ER Triage Study

Overview Of The Research

A groundbreaking study published in Science has examined the performance of large language models in medical diagnostics, including real-life emergency room scenarios. Conducted by a team of physicians and computer scientists from Harvard Medical School and Beth Israel Deaconess Medical Center, the research evaluated how advanced AI models, such as OpenAI’s o1 and 4o, compare to internal medicine physicians in making critical triage decisions.

Methodology And Comparative Analysis

The study analysed cases involving 76 patients treated in the Beth Israel emergency department. Diagnoses made by two internal medicine attending physicians were compared with those generated by the AI models. A separate panel of two blinded attending physicians reviewed all diagnoses to ensure consistency in evaluation. At the triage stage, when patient information was limited, the o1 model matched or exceeded physician accuracy in several cases.

Key Findings And Implications

The o1 model achieved exact or near-exact diagnoses in 67% of cases at triage. In comparison, one physician reached similar accuracy in 55% of cases, while another achieved 50%. Arjun Manrai, head of an AI lab at Harvard Medical School and a lead author of the study, said the model performed above both prior systems and physician baselines.

Limitations And Future Directions

The authors cautioned against allowing AI systems to take on full decision-making roles in life-or-death scenarios at this stage. Experiments were conducted using only text-based data extracted directly from electronic medical records without pre-processing, which limits how broadly the results can be applied. This, in turn, points to the need for further prospective trials in real-world clinical settings. Current models also remain constrained in their ability to process and reason over non-text inputs.

Expert Perspectives And Accountability Concerns

Adam Rodman, a study author, said that the use of AI in clinical settings requires defined accountability frameworks. Emergency physician Kristen Panthagani noted that comparisons with internal medicine physicians, rather than emergency specialists, may affect the interpretation of results. She added that triage decisions focus on identifying potentially life-threatening conditions rather than determining a final diagnosis.

Conclusion

This study emphasizes both the potential and the caution required in integrating AI into critical medical decisions. As the relationship between AI technologies and clinical practice evolves, further rigorous testing and the establishment of accountability frameworks will be indispensable in ensuring that these tools can enhance patient care without compromising safety.

A New Twitter-Inspired Social Network Is Taking Shape

A new social network called Twitter.now is entering the market, with a founding team that includes former Twitter trademark counsel Stephen Coates. The service is being developed by startup Operation Bluebird.

As Ars Technica reported, X sued the company last year and asked a Delaware judge to block the launch. Operation Bluebird argued in a petition that X had abandoned trademarks including “Twitter” and “Tweet.”

Coates has said the project is not an attempt to recreate the original Twitter. In a LinkedIn post, he described the platform as a new public space focused on trust, transparency and user choice.

AI System To Rate Posts

Twitter.now is currently being tested, with early access priced at $20. Its main feature is VERA, an AI system designed to evaluate posts, verify claims and provide sources and context.

Posts receive a trust score, with users eventually able to set a minimum score to filter their feeds. The company says this approach will give people more control over what they see instead of leaving those decisions entirely to an algorithm.

Moderation Remains A Challenge

Scaling moderation will be one of the platform’s biggest tests. Social networks have repeatedly struggled with content moderation as their communities grow, and newer platforms such as Bluesky have faced similar criticism.

Operation Bluebird says VERA will form the basis of its moderation and verification system. A second version is already planned, with expanded tools that would let users set a specific trust threshold for the posts appearing in their feeds.

For now, Twitter.now remains in an early testing phase, combining the familiarity of the Twitter name with an AI-driven approach to evaluating online information.

eCredo
The Future Forbes Realty Global Properties
Uol
Aretilaw firm

Become a Speaker

Become a Speaker

Become a Partner

Subscribe for our weekly newsletter