Enhancing systematic reviews in orthodontics: a comparative examination of GPT-3.5 and GPT-4 for generating PICO-based queries with tailored prompts and configurations-Reference-Cited by-同舟云学术

Enhancing systematic reviews in orthodontics: a comparative examination of GPT-3.5 and GPT-4 for generating PICO-based queries with tailored prompts and configurations

Published:2024-03-07 Issue:2 Volume:46 Page:
ISSN:0141-5387
Container-title:European Journal of Orthodontics
language:en
Short-container-title:

Author:

Demir Gizem Boztaş¹,Süküt Yağızalp¹,Duran Gökhan Serhat¹,Topsakal Kübra Gülnur¹,Görgülü Serkan¹

Affiliation:

1. Department of Orthodontics, Gulhane Faculty of Dentistry, University of Health Sciences , Ankara, Türkiye

Abstract

Summary Objectives The rapid advancement of Large Language Models (LLMs) has prompted an exploration of their efficacy in generating PICO-based (Patient, Intervention, Comparison, Outcome) queries, especially in the field of orthodontics. This study aimed to assess the usability of Large Language Models (LLMs), in aiding systematic review processes, with a specific focus on comparing the performance of ChatGPT 3.5 and ChatGPT 4 using a specialized prompt tailored for orthodontics. Materials/Methods Five databases were perused to curate a sample of 77 systematic reviews and meta-analyses published between 2016 and 2021. Utilizing prompt engineering techniques, the LLMs were directed to formulate PICO questions, Boolean queries, and relevant keywords. The outputs were subsequently evaluated for accuracy and consistency by independent researchers using three-point and six-point Likert scales. Furthermore, the PICO records of 41 studies, which were compatible with the PROSPERO records, were compared with the responses provided by the models. Results ChatGPT 3.5 and 4 showcased a consistent ability to craft PICO-based queries. Statistically significant differences in accuracy were observed in specific categories, with GPT-4 often outperforming GPT-3.5. Limitations The study’s test set might not encapsulate the full range of LLM application scenarios. Emphasis on specific question types may also not reflect the complete capabilities of the models. Conclusions/Implications Both ChatGPT 3.5 and 4 can be pivotal tools for generating PICO-driven queries in orthodontics when optimally configured. However, the precision required in medical research necessitates a judicious and critical evaluation of LLM-generated outputs, advocating for a circumspect integration into scientific investigations.

Publisher

Oxford University Press (OUP)

Link

https://academic.oup.com/ejo/article-pdf/46/2/cjae011/56901599/cjae011.pdf

Reference25 articles.

1. Using artificial intelligence methods for systematic review in health sciences: a systematic review;Blaizot,2022

2. What is the current state of artificial intelligence applications in dentistry and orthodontics;Fawaz;Journal of Stomatology Oral Maxillofacial Surgery,2023

3. Development and accuracy of artificial intelligence-generated prediction of facial changes in orthodontic treatment: a scoping review;Zhu;Journal of Zhejiang University. Science. B.,2023

4. Understanding the capabilities, limitations, and societal impact of large language models;Tamkin,2021

5. Large language models are zero-shot reasoners;Kojima;Advances in Neural Information Processing Systems,2022

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ChatGPT in orthodontics: limitations and possibilities;Australasian Orthodontic Journal;2024-07-01