[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"trial:NCT07457840":3,"trial-entities:NCT07457840":95,"trial-summary:NCT07457840":100},{"id":4,"nct_id":4,"org_study_id":5,"brief_title":6,"official_title":7,"overall_status":8,"completion_date":9,"status_verified_date":10,"last_update_date":11,"start_date":12,"sponsor_name":13,"lead_sponsor_class":14,"has_dmc":15,"brief_summary":16,"detailed_description":17,"conditions":18,"keywords":20,"study_type":21,"primary_purpose":14,"phases":22,"enrollment_info":24,"interventions":27,"primary_outcomes":45,"secondary_outcomes":50,"sex":60,"minimum_age":61,"maximum_age":62,"healthy_volunteers":63,"eligibility_criteria":64,"std_ages":75,"locations":78,"central_contacts":88,"overall_officials":89,"references":93,"see_also_links":94},"NCT07457840","26-45852","Integrating AI Predictions With Clinician Expertise","Transforming Clinical Decision Support Systems: Using Continuous Bayesian Updates to Integrate AI Predictions With Clinician Expertise","ENROLLING_BY_INVITATION","2026-12","2026-07","2026-07-14","2026-02-15","University of California, San Francisco","OTHER",false,"Optimizing the interaction between the human and the machine is a major topic when deploying artificial intelligence (AI) at the bedside. The goal of this randomized clinical vignette study is to learn if presenting AI model outputs via continuous Bayesian updates and\u002For uncertainty quantification can improve diagnostic accuracy and clinician trust in healthcare professionals (physicians, residents, fellows, physician assistants (PAs), and nurse practitioners (NPs)) from US academic institutions evaluating patients with chest pain or dyspnea.\n\nThe main questions it aims to answer are:\n\n* Does presenting AI predictions as Bayesian-updated post-test probabilities improve diagnostic accuracy compared to standard predicted probabilities?\n* Does the addition of uncertainty quantification (95% confidence intervals) to AI predictions improve diagnostic accuracy?\n* Do these interventions (Bayesian updating and\u002For uncertainty quantification) help clinicians recover from the negative effects of intentionally misleading AI predictions?\n\nComparison: Researchers will compare standard AI predicted probabilities (presented without uncertainty) to Bayesian-updated post-test probabilities and\u002For outputs containing 95% confidence intervals to see if the interventions improve diagnostic accuracy, clinician confidence, and resilience against misleading AI.\n\nParticipants will:\n\n* Review 8 clinical vignettes (simulated patient cases) focusing on chest pain or dyspnea.\n* Provide an initial \"pre-test\" diagnostic probability for 5 possible diagnoses based on the clinical history alone.\n* View AI model outputs that vary by experimental condition (standard probability vs. Bayesian update, with or without uncertainty intervals, and accurate vs. misleading).\n* Provide an updated \"post-test\" diagnostic probability for the diagnoses after viewing the AI output.\n* Select and rank diagnostic tests and therapeutic steps for each vignette. Complete a post-survey regarding their trust in the AI, comfort with the data presentation, and demographics.","Study Design: This is a 2x2 factorial within-subjects design. The two factors are (1) Bayesian updating via continuous likelihood ratios (CLR) vs. standard predicted probability, and (2) uncertainty quantification (95% confidence intervals) vs. point estimate only. AI prediction accuracy (accurate vs. intentionally misleading) is varied as a within-subjects stratification factor balanced across all 4 conditions, with half of each participant's vignettes receiving accurate predictions and half receiving misleading predictions. AI predictions are simulated (pre-programmed) for experimental control. Vignette order and condition assignment are independently randomized per participant.\n\nPrimary Analysis: Diagnostic accuracy is analyzed using a generalized linear mixed model (GLMM) with fixed effects for CLR, Uncertainty, Misleading, and vignette, and a participant random intercept. Pre-specified secondary analyses examine interactions of presentation format with misleading AI.\n\nSample Size: Simulation-based power analysis (1,000 Monte Carlo iterations per scenario) was conducted using the planned GLMM. Assuming 70% baseline diagnostic accuracy and within-participant ICC of 0.25, the study achieves 85.8% power for the CLR main effect and 85.7% for the Uncertainty main effect with N=100 at alpha=0.05 (two-tailed).",[19],"Diagnostic Decision Making",[],"INTERVENTIONAL",[23],"NA",{"count":25,"type":26},100,"ESTIMATED",[28,35,41],{"type":29,"name":30,"description":31,"armGroupLabels":32},"BEHAVIORAL","Bayesian-Updated Post-Test Probability","Rather than presenting the AI model's raw predicted probability, the system takes the clinician's pre-test probability (entered before seeing AI output) and applies a continuous likelihood ratio (CLR) derived from the AI model to calculate a Bayesian-updated post-test probability. The output is displayed as a shift from the clinician's own assessment (e.g., \"Your assessment: 45% -\\> Updated assessment: 72%\"). The raw AI prediction is not shown. This approach mirrors how clinicians use diagnostic test results such as D-dimer to update pre-test probability of pulmonary embolism.",[33,34],"Bayesian Updating (CLR) + No Uncertainty","Bayesian Updating (CLR) + Uncertainty (95% CI)",{"type":29,"name":36,"description":37,"armGroupLabels":38},"Standard AI Predicted Probability","AI model prediction is presented as a simple predicted probability (0-100%) for each of the possible diagnoses, together with the top 3 clinical features driving the prediction (e.g., \"Acute Myocardial Infarction: 68% - Key factors: elevated troponin, ST-segment changes on ECG, chest pain radiation to left arm\"). This represents the most common current approach to presenting AI-based diagnostic predictions in clinical settings.",[39,40],"Standard Probability + No Uncertainty (Control)","Standard Probability + Uncertainty (95% CI)",{"type":29,"name":42,"description":43,"armGroupLabels":44},"Uncertainty Quantification (95% Confidence Interval)","The AI output (whether Bayesian-updated post-test probability or standard predicted probability) is presented together with a 95% confidence band displayed as error bars on probability bars. For accurate AI predictions, confidence interval width is approximately +\u002F-12-15 percentage points. For misleading AI predictions, confidence intervals are widened by a factor of 1.5x (approximately +\u002F-18-23 percentage points) to simulate reduced model confidence in unfamiliar or edge-case scenarios. Confidence intervals are constrained to the 0-100% range.",[34,40],[46],{"measure":47,"description":48,"timeFrame":49},"Clinician Diagnostic Accuracy","Proportion of correct diagnostic assessments across all vignettes and experimental conditions. For each vignette, participants rate 5 possible diagnoses on a 0-100% probability scale. The diagnosis assigned the highest probability is considered the participant's final diagnosis. Accuracy is determined by comparing the final diagnosis to the ground truth diagnosis established by expert panel consensus (minimum 4 of 5 board-certified physicians in agreement). Analyzed using a generalized linear mixed model (GLMM) with binary outcome (correct vs. incorrect), fixed effects for CLR, uncertainty quantification, misleading AI, and vignette, and a random intercept for participant.","Day 1 during survey completion",[51,54,57],{"measure":52,"description":53,"timeFrame":49},"Change in Diagnostic Probability Estimates","Magnitude and direction of change in clinician-provided probability estimates from pre-test assessment (before AI output) to post-test assessment (after AI output) for each of 5 possible diagnoses per vignette. Measured on a 0-100% scale.",{"measure":55,"description":56,"timeFrame":49},"Diagnostic Accuracy Under Misleading AI Predictions","Proportion of correct final diagnoses when AI predictions are intentionally misleading vs. accurate, and whether the interventions (Bayesian updating, uncertainty quantification) mitigate the negative effect of misleading AI. Assessed via interaction terms (CLR x Misleading, Uncertainty x Misleading) in the primary GLMM.",{"measure":58,"description":59,"timeFrame":49},"Clinician Satisfaction With AI Decision Support (Exploratory)","Self-reported satisfaction with the AI-based clinical decision support, measured via question(s) in the post-survey questionnaire.","ALL","18 Years",null,true,{"inclusion":65,"exclusion":69,"raw_text":74},[66,67,68],"Must hold one of the following clinical roles: Nurse Practitioner (NP), Physician Assistant\u002FPhysician Associate (PA), Resident Physician, Physician Fellow, or Attending Physician","Able to complete the survey in English","Access to a computer or tablet (mobile phones are not recommended due to the visual nature of the survey)",[70,71,72,73],"Does not hold an eligible clinical role as defined above","Completes fewer than 2 of 8 clinical vignettes (less than 25% of the survey)","Has previously participated in this study","Unable to complete the survey in English","Inclusion Criteria:\n\n* Must hold one of the following clinical roles: Nurse Practitioner (NP), Physician Assistant\u002FPhysician Associate (PA), Resident Physician, Physician Fellow, or Attending Physician\n* Able to complete the survey in English\n* Access to a computer or tablet (mobile phones are not recommended due to the visual nature of the survey)\n\nExclusion Criteria:\n\n* Does not hold an eligible clinical role as defined above\n* Completes fewer than 2 of 8 clinical vignettes (less than 25% of the survey)\n* Has previously participated in this study\n* Unable to complete the survey in English",[76,77],"ADULT","OLDER_ADULT",[79],{"facility":80,"city":81,"state":82,"zip":83,"country":84,"geoPoint":85},"ZSFG","San Francisco","California","94110","United States",{"lat":86,"lon":87},37.77493,-122.41942,[],[90],{"name":91,"affiliation":13,"role":92},"Romain Pirracchio, MD, PhD, MPH","PRINCIPAL_INVESTIGATOR",[],[],{"nct_id":4,"conditions":96,"biomarkers":99},[97,19,98],"Chest Pain","Dyspnea",[],{"nct_id":4,"found":63,"summary":101,"prompt_version":111},{"design":102,"status":103,"heading":104,"summary":105,"follow_up":106,"word_count":107,"commitments":108,"compensation":109,"drugs_mentioned":110},"This study is an interventional study with 100 participants. It uses a 2x2 factorial within-subjects design, meaning each participant will experience different ways of seeing AI predictions.","completed","Improving Diagnostic Accuracy with AI Tools","This study is looking at how different ways of showing artificial intelligence (AI) predictions can help doctors and other healthcare professionals make more accurate diagnoses. It's testing two main approaches: one called \"Bayesian-Updated Post-Test Probability,\" which combines a clinician's initial thoughts with AI information, and another called \"Standard AI Predicted Probability,\" which just shows the AI's prediction. The study also looks at whether showing a \"95% Confidence Interval\" (a range that shows how sure the AI is) helps. The goal is to see if these methods improve diagnostic accuracy for patients with chest pain or shortness of breath. We are looking for 100 healthcare professionals, including Nurse Practitioners, Physician Assistants, and Doctors, who can complete an online survey in English.","Participants are followed for diagnostic accuracy, which is measured on Day 1 during survey completion.",120,"Participants will complete an online survey using a computer or tablet. The primary endpoint, clinician diagnostic accuracy, is measured on Day 1 during survey completion.","Not stated in the trial record.",[30,36],"v2"]