MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice

Rodolfo Rizzi1 · Alessandro Grecucci2,† · Massimo Stella1,†

1 CogNosco Lab, Department of Psychology and Cognitive Science, University of Trento, Italy 2 Department of Education, Psychology and Communication Sciences, University of Bari, Italy † Co-last authors† Ultimi autori a pari merito

Abstract

AbstractAbstract (originale in inglese)

Psychotherapists need repeated training and supervision; however, scalability is problematic. We present MyMentorLLM, a multimodal voice- and text-based deliberate-practice environment with 2,100 complete Cognitive Behavioural Therapy (CBT) sessions. Each session links a DSM-5-TR-grounded LLM patient (with major depressive, generalised anxiety or borderline personality disorder), an LLM therapist-in-training and an LLM expert supervisor (powered by Gemma-4, Gemini-3.1-Flash-Live and Qwen-3.6).

Sessions were analysed for emotional dynamics, therapeutic competence and diagnostic accuracy against human psychotherapy data. Simulated patients expressed disorder-congruent emotional profiles, which therapists mirrored as in human counselling. LLM trainee competence was rated above human levels in most conditions, while native speech-to-speech was closest to human scores. Supervisor feedback improved diagnostic accuracy in 5 of 7 LLM conditions, whereas symptom identification accuracy increased with model size. This work shows deliberate practice can be simulated for CBT training, although patient fidelity, supervisor calibration and harmful feedback require evaluation via a complex systems perspective.

SessionsSedute
2,100
Dialogue turnsTurni di dialogo
65,100
Spoken therapy, 900 audio sessionsParlato, 900 sedute audio
263.4 h
Audio clips (FLAC)Clip audio (FLAC)
27,900 · 20.3 GiB
LicenceLicenza
CC BY 4.0

The simulated mentorship cycleIl ciclo simulato di supervisione

Figure 1 of the paper. Panel a: three stages, setup (model backends, patient disorder personas, fixed trainee and mentor), a first CBT session between patient and trainee in voice or text for 31 turns, and supervision with CTRS scores, a reflective question, diagnosis asked again and symptom recognition. Panel b: the persona system prompts for patient, trainee therapist and CBT mentor.
Fig. 1. a. The simulated CBT mentorship cycle. A simulated session consists of three stages: (1) setup, where the language models and a patient persona grounded in DSM-5-TR clinical cases are selected while the trainee therapist and mentor remain fixed; (2) an autonomous voice- or text-based CBT session between the patient and trainee, with audio and transcripts recorded; and (3) mentor supervision, including CTRS scoring, diagnostic feedback, and DSM-5-TR symptom recognition. b. Persona system prompts. The cards show the common structure of each persona's prompt, comprising a persona definition, a task, and role-specific instruction blocks.

What we foundCosa abbiamo trovato

Across 2,100 sessions we analysed how the simulated patients express their disorder, how the trainees' competences are rated, and what the supervisor's feedback actually changes.Su 2.100 sedute abbiamo analizzato come i pazienti simulati esprimono il disturbo, come vengono valutate le competenze dei terapeuti e che effetto ha davvero il feedback del supervisore.

Simulated patients express disorder-congruent emotional profiles. Therapists mirror them, as in human counselling.I pazienti simulati esprimono profili emotivi coerenti con il disturbo. I terapeuti li rispecchiano, come nel counselling umano.

Each patient is grounded in a DSM-5-TR clinical case. Across conditions the simulated patients showed disorder-congruent emotional profiles: depression was marked by sadness, generalised anxiety by fear and anticipation, and borderline personality disorder by broad negative affect, including fear, sadness and anger. Therapists mirrored the patient’s affect in attenuated form while maintaining trust and anticipation, a pattern also observed in the human HOPE counselling corpus.Ogni paziente è basato su un caso clinico del DSM-5-TR. Nelle diverse condizioni i pazienti simulati mostrano profili emotivi coerenti con il disturbo: la depressione è caratterizzata dalla tristezza, l’ansia generalizzata da paura e anticipazione e il disturbo borderline da un più ampio spettro di emozioni negative, tra cui paura, tristezza e rabbia. I terapeuti rispecchiano l’affetto del paziente in forma attenuata, mantenendo fiducia e anticipazione, un pattern osservato anche nel corpus umano HOPE.

Fig. 2 · Supplementary Table 1

Most AI supervisors overrate the AI therapist-in-training. The native speech-to-speech model comes closest to human scores.Quasi tutti i supervisori AI sopravvalutano il terapeuta AI in formazione. Il modello nativo speech-to-speech è il più vicino ai punteggi umani.

In the Gemini 3.1 Live · LA condition patient and therapist speak and listen over an audio channel, as in a real conversation, so pacing, prosody and hesitation are real. Those sessions are the ones scored closest to therapy by real trainees: 30.0 on average, against 31.0 for 1,264 sessions by human community clinicians. Where answers are written and voiced afterwards, or stay written, scores climb from 41.2 to 56.0.Nella condizione Gemini 3.1 Live · LA paziente e terapeuta si parlano e si ascoltano su un canale audio, come in una conversazione vera, quindi ritmo, prosodia ed esitazioni sono reali. Sono le sedute con i punteggi più vicini a quelli della terapia condotta da terapeuti in formazione veri: 30,0 in media, contro 31,0 di 1.264 sedute di clinici umani. Dove le risposte sono scritte e poi lette da una voce sintetica, o restano scritte, i punteggi salgono da 41,2 a 56,0.

Table 1 · Fig. 3

Mentor feedback improved diagnoses in 5 of 7 conditions.Il feedback del mentor ha migliorato le diagnosi in 5 condizioni su 7.

The improvement came mainly from correcting initially incorrect BPD diagnoses. Overall, BPD accuracy rose from 65.3% to 82.6%.Il miglioramento riguarda soprattutto la correzione delle diagnosi BPD inizialmente errate. Nel complesso, l'accuratezza sul BPD sale dal 65,3% all'82,6%.

Fig. 4

Symptom recognition improves with model size.Il riconoscimento dei sintomi cresce con la dimensione del modello.

Symptom-identification accuracy increases with model size: Qwen 3.6 35B reaches 95.5%, Gemma 4 12B and Gemini 3.1 Live 86–87%, while the smaller Gemma 4 E2B falls to 19–26%.L'accuratezza nell'identificazione dei sintomi aumenta con la dimensione del modello: Qwen 3.6 35B raggiunge il 95,5%, Gemma 4 12B e Gemini 3.1 Live l'86–87%, mentre il più piccolo Gemma 4 E2B scende al 19–26%.

Fig. 4a · Fig. 5

Feedback can also do harm. Fidelity, calibration and safety must be judged together.Il feedback può anche nuocere. Fedeltà, calibrazione e sicurezza vanno valutate insieme.

With the smallest model, Gemma 4 E2B, 12–16% of correct diagnoses were abandoned after the mentor spoke, against 4.8% for Qwen 3.5 9B and virtually none for the larger Gemma 4 12B, Gemini 3.1 Live and Qwen 3.6 35B. For weaker trainees, supervision became a trigger to defer.Con il modello più piccolo, Gemma 4 E2B, il 12–16% delle diagnosi corrette è stato abbandonato dopo l'intervento del mentor, contro il 4,8% di Qwen 3.5 9B e quasi nessuna per i più grandi Gemma 4 12B, Gemini 3.1 Live e Qwen 3.6 35B. Per i terapeuti più deboli, la supervisione diventa una spinta ad adeguarsi.

Results and limitationsRisultati e limiti

In the charts: LA = native live audio, A = written answers voiced afterwards by a separate voice model (OmniVoice), T = text only.Nei grafici: LA = audio nativo dal vivo, A = risposte scritte e poi lette da un modello vocale separato (OmniVoice), T = solo testo.

The datasetIl dataset

Dataset and codebase are public:Dataset e codice sono pubblici: Dataset on Hugging Face, CC BY 4.0Dataset su Hugging Face, CC BY 4.0 Analysis code on GitHub, MITCodice delle analisi su GitHub, MIT

All sessions are available in the open dataset, with turn-by-turn transcripts and structured mentor evaluations; for 900 sessions, audio is also available for every spoken turn.Tutte le sedute sono disponibili nel dataset open, con trascrizioni turno per turno e valutazioni strutturate del mentor; per 900 sedute è disponibile anche l’audio di ogni turno parlato.

Quick startPer iniziare

import pandas as pd

url = "hf://datasets/RodolfoRizzi/MyMentorLLM-dataset/data/"
turns    = pd.read_parquet(url + "therapy_turns.parquet")     # 65100 x 11
sessions = pd.read_parquet(url + "mentor_sessions.parquet")   # 2100 x 46

Content noticeAvvertenza sui contenuti

The conversations are model-generated: no real patients took part and the files contain no personal data. The patient personas are clinically grounded, inspired by the published vignettes of DSM-5-TR Clinical Cases (Barnhill, 2023) and the DSM-5-TR criteria, so the transcripts contain realistic clinical material, including expressions of suicidal ideation and self-harm.Le conversazioni sono generate da modelli: nessun paziente reale ha partecipato e i file non contengono dati personali. I pazienti sono clinicamente fondati, ispirati alle vignette pubblicate di DSM-5-TR Clinical Cases (Barnhill, 2023) e ai criteri del DSM-5-TR, quindi le trascrizioni contengono materiale clinico realistico, incluse espressioni di ideazione suicidaria e autolesionismo.

CiteCome citare

@article{rizzi2026mymentorllm,
  title   = {MyMentorLLM: A psychotherapy GenAI environment with multimodal
             voice/text patients, trainees and experts for deliberate practice},
  author  = {Rizzi, Rodolfo and Grecucci, Alessandro and Stella, Massimo},
  journal = {arXiv preprint arXiv:2607.25667},
  year    = {2026}
}

CITATION.cff · LicenceLicenza CC BY 4.0