Society for the Advancement of Psychotherapy

Artificial Intelligence and Psychotherapy: Opportunities, Challenges, and Recommendations

Society for the Advancement of Psychotherapy

Society for the Advancement of Psychotherapy

September 4, 2026

Artificial Intelligence and Psychotherapy: Opportunities, Challenges, and Recommendations

A Report of the AI and Psychotherapy Work Group Society for the Advancement of Psychotherapy American Psychological Association, Division 29

September 2026

Work Group Chair:

Stewart Cooper, PhD

Work Group Members:

Firouz Ardalan, PhD

John Gavazzi, PsyD 

Astrea Greig, PsyD

Gerry Koocher, PhD         

Erica Lee, PhD 

Xu Li, PhD 

James Lichtenberg, PhD                                                               

Wonjin Sim, PhD   

Melanie Wilcox, PhD

Executive Summary

Psychologists are already using artificial intelligence (AI) to draft progress notes, screen literature, generate case conceptualizations, and rehearse clinical skills with simulated patients. Technology has shaped professional psychology for decades, but no earlier development moved this quickly into so many parts of the work at once. A model can be given every word a patient has spoken and still have nothing at stake in what happens next. That disparity organizes the analysis in this report.

Emerging from the October 2025 SAP Board discussion of mega-level issues and opportunities for the Society for the Advancement of Psychotherapy, the Presidential Work Group on AI and Psychotherapy was established in 2026 by Society President Joshua Swift to examine AI’s implications for psychotherapy practice, supervision, education and training, and research. What follows is guidance from a Society work group. It is not APA policy and not a practice guideline.

The authors of the four domain sections worked independently and drew on different literatures. Their analyses converged: AI can substantially augment psychological work, but it cannot replace the human judgment, relationships, and accountability on which psychotherapy depends.

Six principles follow.

  1. AI augments rather than replaces humans.
  2. Human accountability remains essential.
  3. AI output requires critical evaluation.
  4. Implementation must be ethical and safe.
  5. Human relationships remain central.
  6. Continuous learning and governance are necessary.

Across all four domains the Work Group identified the same risks: automation bias, professional deskilling, cultural and demographic bias, threats to privacy and confidentiality, and overreliance on AI-generated recommendations. Beneath them sits a single concern. AI generates plausible interpretations from patterns in data, while psychotherapy requires understanding a particular person within a particular cultural, historical, and relational context. Pattern recognition can look like clinical understanding and be something else entirely.

The Work Group recommends evaluating any AI application against three questions: whether it enhances or erodes the human relationships the work depends on, whether it preserves psychologist accountability, and whether it respects the patient’s narrative integrity and cultural context. Evaluate a tool before it becomes embedded in practice, because a tool is far easier to decline than to remove.

Four practices follow. Psychologists should reason independently before consulting AI models whenever feasible, verify what AI produces, protect confidentiality in data handling, and review AI-assisted documentation before it enters a record.

Underlying these recommendations is a commitment to human dignity. Persons must not be reduced to data, classifications, or algorithmic predictions. The tools named in this report will be superseded and the studies cited will be replaced by better ones, but the question they raise will not date: what parts of psychological work can be augmented by technology, and what responsibilities cannot be delegated at all.

Report of the Society for the Advancement of Psychotherapy

Work Group on AI and Psychotherapy

Stewart Cooper and John Gavazzi

In October 2025, then-President Stewart Cooper of the Society for the Advancement of Psychotherapy (APA Division 29) led the Society’s Executive Board in a discussion of the opportunities, challenges, risks, and benefits that artificial intelligence (AI) presents for psychotherapy, including its implications for clinical practice, supervision, education and training, and research. The Board unanimously agreed that AI represented a strategic issue of critical importance to the profession and that the Society should provide leadership and resources to help its members understand and respond to these rapidly evolving developments. Building on that discussion, 2026 Society President Joshua Swift established a Presidential Work Group to examine the implications of AI for psychotherapy and to develop guidance and resources for the profession.

The report that follows is the product of the AI and Psychotherapy Work Group’s efforts. It is organized into six sections. The report begins with a brief introduction about AI and psychotherapy including six core principles, followed by four sections respectively addressing AI’s implications for psychotherapy practice, supervision, education & training, and research. The final segment offers overarching reflections and recommendations. Regarding the body of the report, AI and Psychotherapy Practice was authored by Astrea Greig and Erica Lee. AI and Psychotherapy Supervision was written by John Gavazzi, Gerry Koocher, and Melanie Wilcox. AI and Psychotherapy Education and Training was prepared by Firouz Ardalan and James Lichtenberg. AI and Psychotherapy Research was authored by Wonjin Sim and Xu Li. Stewart Cooper, Chair of the Presidential Work Group, and John Gavazzi co-authored the Introduction and the concluding Reflections and Recommendations.  

Introduction

We are entering a new era marked by both extraordinary opportunity and unprecedented uncertainty, driven in large part by the remarkable pace, scope, and transformative potential of technological change (Kissinger et al., 2021). Technology has influenced the practice, supervision, education and training, and research of psychotherapy for several decades (American Psychological Association [APA], 2013; Maheu et al., 2017; Norcross & Lambert, 2019). No previous technological innovation, however, appears poised to rival the breadth, depth, and speed of transformation that AI is expected to bring to the practice, science, education, and supervision of psychology (APA, 2025; Topol, 2019; World Economic Forum, 2025).

Although the four sections in this consolidated article address different domains of professional psychotherapy, they converge on a common set of themes that can be distilled into six overarching principles. Together, these principles provide a framework to guide the profession as it navigates the opportunities and challenges that AI presents.   

Principle 1: AI Augments Rather Than Replaces Humans      

AI tools should extend the reach of psychologists, supervisors, educators, and researchers rather than substitute for them. The reason is not that the technology is immature but that the capacities on which this work depends are constitutive of being human. Human dignity is not earned, not algorithmically assessed, and not reducible to data. Any application that treats a patient as a pattern to be classified rather than a person to be encountered is not merely technically limited but morally deficient. The empathy, attachment, and cooperative moral sense that make psychotherapy possible are products of biological and cultural evolution, matured in vivo across a life of actual relationships (Gavazzi, 2025). AI systems operate on approximations of those experiences distilled from training data, and they lack a robust theory of mind, meaning the capacity to attribute beliefs, intentions, and emotions to oneself and to another. A model can be given every word a patient has spoken and still have nothing at stake in what happens next. The psychologist has a great deal at stake, and that asymmetry is not a limitation awaiting a better system. It is the condition under which the work becomes possible at all.

Principle 2: Human Accountability     

Final responsibility for clinical, supervisory, educational, and scientific decisions remains with the qualified human professional, regardless of what informed the decision. The Ethics Code (APA, 2017) assigns that accountability unambiguously, and APA’s (2025) guidance on AI in health service psychology affirms that human oversight is not diminished by the sophistication of the tool. Accountability is also more than a procedural assignment; it is exercised through the practitioner’s own sense of responsibility for a judgment, and that sense is precisely what habitual reliance on AI output tends to erode. A system that generates a diagnosis, risk determination, or competency rating without a human in the loop is not functioning as a clinical tool. As Wiener (1948) recognized in the earliest days of cybernetics, technological capability neither replaces nor diminishes the irreducibly human responsibility at the center of meaningful human processes, and nothing in the current generation of these systems changes that.    

Principle 3: Critical Evaluation      

AI outputs require verification, informed skepticism, and independent professional judgment. Contemporary large language models generate text through probabilistic pattern recognition, producing responses that are coherent and contextually responsive while possessing no internal understanding of the material (Bender et al., 2021). Fluency is therefore not evidence of soundness. Fabricated citations, plausible but contraindicated recommendations, and confident accounts of law or regulation that omit controlling authority are characteristic failure modes rather than rare malfunctions. Two implications follow. First, AI-generated material is best treated as an object of critical analysis rather than as a draft awaiting light editing. Second, verification is bounded by what the practitioner already knows, because an omission is invisible to a psychologist who cannot name what is missing. Awareness alone is an insufficient safeguard. Physicians who had completed dedicated AI literacy training remained susceptible to subtly erroneous recommendations (Qazi et al., 2026). Safeguards must therefore be structural, beginning with the practice of reasoning independently before consulting any AI model.

Principle 4: Ethical and Safe Implementation 

Privacy, informed consent, transparency, and fairness are preconditions for responsible use rather than considerations to be addressed afterward. Any tool that accesses, processes, or stores protected health information presupposes HIPAA and HITECH obligations, including with a signed business associate agreement, and many widely used general-purpose platforms explicitly disclaim compliance. De-identification is the minimum standard, and it is harder to achieve than it appears since the details that make a case specific enough to analyze are often the details that make the parties identifiable. Patients and supervisees are entitled to meaningful disclosure of how these tools are used in their care or training, communicated in a culturally and linguistically appropriate manner, along with a genuine opportunity to decline. Fairness deserves equal weight. AI systems encode and can amplify existing inequities, whether through the design choices built into a system (Obermeyer et al., 2019) or through disparities in the outputs themselves (Omar et al., 2025), and appraising outputs for embedded cultural and discriminatory assumptions is a professional obligation rather than an optional refinement.

Principle 5: Human Relationships Remain Central 

The therapeutic, supervisory, educational, and scientific relationship is the mechanism through which this work succeeds, and it cannot be delegated. Decades of psychotherapy research identify the quality of the therapeutic alliance as among the strongest predictors of outcome across theoretical orientations and clinical populations. Similarly, the supervisory alliance has been described as the heart of supervision for more than five decades. What constitutes a genuine alliance, however, is culturally contextual. Presence may mean sustained eye contact in one cultural frame and attentive listening without constant eye gaze in another, and no generic model of therapeutic presence can supply that distinction. A language model has no access to the relationship, the moment, or the situated cultural meaning required to interpret it. The relevant contrast is between abstract synthesis, assembled from symbolic representations of clinical reality, and concrete synthesis, which is embodied and embedded in a particular relationship.

Principle 6: Continuous Learning and Governance 

Ongoing research, AI literacy, policy development, and ethical oversight are necessary as these technologies evolve. The empirical base remains uneven, and the regulatory framework, including the Ethics Code, HIPAA, HITECH, and FERPA, predates the technology it is now being asked to govern. Practitioners, supervisors, and training programs need not wait for definitive standards in order to act responsibly. Establishing a written policy on AI use, verifying vendor agreements and meaningful consent, teaching critical appraisal and prompt construction explicitly, and treating AI-generated data as supplementary rather than standalone are all available now. Where research is thin and existing frameworks leave gaps, the principles of autonomy, beneficence, nonmaleficence, fidelity, and justice provide a stable basis for defensible decisions. Governance also has a disciplinary dimension. Psychology has a substantial interest in shaping how these tools are developed, validated, and deployed, and that influence depends on sustained engagement rather than on either uncritical adoption or categorical refusal. 

References: Introduction

American Psychological Association. (2013). Guidelines for the practice of telepsychology. American Psychologist, 68(9), 791–800. https://doi.org/10.1037/a0035001

American Psychological Association. (2017). Ethical principles of psychologists and code of conduct (2002, amended effective June 1, 2010, and January 1, 2017). https://www.apa.org/ethics/code

American Psychological Association. (2025). Ethical guidance for AI in the professional practice of health service psychology. https://www.apa.org/topics/artificial-intelligence-machine-learning/ethical-guidance-professional-practice.pdf

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 610–623). Association for Computing Machinery. https://doi.org/10.1145/3442188.3445922

Gavazzi, J. D. (2025, December). Why artificial intelligence will not replace human psychologists: Legal, ethical, and clinical limitations. Psychotherapy Bulletin, 61(1). https://societyforpsychotherapy.org/why-artificial-intelligence-will-not-replace-human-psychologists-legal-ethical-and-clinical-limitations/

Kissinger, H., Schmidt, E., & Huttenlocher, D. (2021). The age of AI: And our human future. Little, Brown and Company.

Maheu, M. M., Drude, K. P., & Wright, S. D. (Eds.). (2017). Career paths in telemental health. Springer.

Norcross, J. C., & Lambert, M. J. (2019). Evidence-based psychotherapy relationships: The third task force. In J. C. Norcross & M. J. Lambert (Eds.), Psychotherapy relationships that work: Evidence-based therapist contributions (3rd ed., pp. 1–23). Oxford University Press. https://doi.org/10.1093/med-psych/9780190843953.003.0001

Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342

Omar, M., Soffer, S., Agbareia, R., Bragazzi, N. L., Apakama, D. U., Horowitz, C. R., Charney, A. W., Freeman, R., Kummer, B., Glicksberg, B. S., Nadkarni, G. N., & Klang, E. (2025). Sociodemographic biases in medical decision making by large language models. Nature Medicine, 31(6), 1873–1881. https://doi.org/10.1038/s41591-025-03626-6

Qazi, I. A., Ali, A., Khawaja, A. U., Akhtar, M. J., Sheikh, A. Z., & Alizai, M. H. (2026). Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial. NEJM AI, 3(5). https://doi.org/10.1056/aioa2501001

Topol, E. (2019). Deep medicine: How artificial intelligence can make healthcare human again. Basic Books.

Wiener, N. (1948). Cybernetics: Or control and communication in the animal and the machine. Technology Press.

World Economic Forum. (2025). The future of jobs report 2025. World Economic Forum.

Artificial Intelligence (AI) Technologies in Clinical Practice

Astrea Grieg and Erica Lee

AI has become a relatively standard part of life within current society, and psychology and mental health care are no exception. Both clients and/or mental health providers may be regular utilizers of AI in their daily activities. A recent study found one in five adolescents use AI for mental health reasons, and a large amount of youth use it monthly or more (McBain et al., 2026). As with any new technology, AI, while incredibly useful, is not without its challenges. Therapists and consumers of mental health care alike must navigate this new frontier in a way that maximizes benefit and reduces harm. Unfortunately, this may entail technological savviness that many do not yet possess and/or require mental health care services to be delivered with policies in place to ensure client safety and privacy. There are a multitude of concerns regarding the use of AI in many aspects of clinical practice. These concerns can be grouped into those regarding the mental health care profession, psychologists and other mental health providers, clients, and psychological intervention concerns. In this section, when referring to AI, we particularly focus on generative AI (also known as Gen AI), which includes large language models (LLMs). LLMs are typically large Gen AI models that can produce text responses to a user’s prompts based on prediction from having been trained on enormous amounts of training data (Hua et al., 2025). Gen AI and LLMs provide more advanced and nuanced responses rather than older decision tree style logic.

AI Concerns Related to Psychology and Mental Health Care

There are numerous regulatory and ethical issues to consider regarding the use of AI in psychological practice. AI tools for mental health care are not regulated or trained in the same manner as psychologists and other mental health providers are. Additionally, many people use general AI models —which are not created to address mental health— for mental health needs, which there is not adequate regulation and guidance for (McBain et al., 2026). Unregulated AI chatbots can pose significant risks, as they can produce inaccurate advice, privacy violations, and even the potential to cause harm (APA, 2025; Augustin et al., 2026; Bhattacharyya et al 2023; Hua et al., 2025; Iftikhar et al, 2025; Torous et al., 2025). There are growing calls for regulation to ensure that AI tools are developed with the input of mental health professionals and include safeguards for users—especially for vulnerable populations (Iftikhar et al, 2025). Mental health providers should warn clients of the pitfalls of AI use for mental health advice (APA, 2025a; Iftikhar et al, 2025).

Likewise, mental health providers should advocate for policymakers to prohibit AI from representing itself as a qualified licensed mental health professional (APA, 2025a; Iftikhar et al., 2025). There is a strong interest in the development of AI therapists due to it seeming feasible on a surface level and likely for financial outcomes. Mental health providers and psychotherapists in particular use clinical discourse and interaction as their clinical instrument of care, and AI chatbots are programmed to be capable of holding conversations with human users. However, AI tools are unfamiliar with clinical treatment protocols and guidelines, and lack experience, training, and supervision. Even if AI becomes familiar with psychotherapy or psychological testing concepts and protocols, they lack practice and understanding of how to clinically apply protocols and interpret results or outcomes.

While in use, AI tools are not referencing the APA ethics code or federal and local laws (Hua et al., 2025). AI is known to make the user want to continue to use it; as such there are concerns about AI sycophancy and AI hallucinations (Augustin et al., 2026; Bhattacharyya et al 2023). AI Sycophancy refers to the tendency for AI responses to be excessively validating and agreeable (Augustin et al., 2026) and AI hallucinations refer to fabricated misinformation (Bhattacharyya et al 2023). AI errors or hallucinations aside, psychotherapy does not typically involve endless validation and agreeableness. Psychotherapists, on the other hand, have a nuanced understanding of the client’s needs and are expected to be adaptively responsive to clients to meet treatment goals. AI additionally does not have a concept of informed consent which is a client’s need to provide consent and understanding of a psychological service or treatment before starting (Saeidnia et. al., 2024, Bloch- Atefi, (2025).

AI cannot fully replace the essential human elements of psychotherapy.The therapeutic relationship, or therapeutic alliance, is a crucial component of effective treatment (Martin et al., 2000) and is something that AI cannot adequately replicate. AI can mimic empathy through language processing but does not possess genuine emotional understanding (Wang et al., 2025). A human therapist’s ability to interpret nuanced cues like body language, tone of voice, and facial expressions, as well as to share a sense of genuine human connection and presence, is critical for building trust and a safe therapeutic space. AI therefore cannot be truly empathic or understanding; and, it lacks clinical intuition (APA, 2025). 

AI systems are based on algorithms and data patterns, which cannot account for the “gray areas” or the nuances of human experience. AI will struggle with ethical dilemmas, complex trauma, and crisis situations, and likely not have the capacity to recognize and intervene when a person is at risk of self-harm. AI tools are not as skilled in managing clinical safety concerns as psychotherapists and other mental health providers are. In crisis moments, mental health providers have to do a number of tasks that AI cannot. This often includes communicating with local emergency departments and/or local EMS, and completing required forms as necessary per the state or region the client is in, all while following local and national healthcare confidentiality laws and keeping the client’s safety in mind. If AI attempts to develop this capacity to facilitate a risk management plan, it would likely struggle and be incapable of carrying out these tasks that are to be completed by not just humans but licensed medical professionals. In total, if using AI, it should be used as a tool to aid in clinical practice but never to replace qualified human mental health professionals (APA, 2025; McBain et al., 2025).

Concerns Related to the Psychologist and Other Mental Health Providers

Currently, APA lacks formal guidelines for AI use, there is only a recent brief advisory and guidance document (APA, 2025; APA, 2025a). Additionally, mental health providers need to abide by federal privacy law (i.e. Health Insurance Portability and Accountability Act [HIPAA]) as well as state or regional laws and organizational policies regarding the storage of protected health information (PHI). With the addition of AI use, mental health providers need to ensure that AI tools used in a clinical setting are compatible with these federal and state laws and organizational policies. Often, at the organizational level, both legal and information technology staff must vet AI tools for use within healthcare visits, including psychotherapy visits, psychological assessment, and other psychological clinical practice, to ensure patient safety and privacy (APA, 2025b). Often, the speed of the development of AI tools outpaces the development of policy and regulation.  

Aside from ascertaining whether using an AI tool is permitted, mental health providers need training on how to properly use AI tools within clinical practice. Individual practitioners, group practices, and organizations that deliver psychological services are responsible for ensuring mental health providers use AI effectively and know how to monitor AI tools with patient safety in mind (APA, 2025). This need for human oversight also increases mental health care provider workload burden and responsibility (Auf et al., 2025; Ni & Jia, 2025). Some AI tools, such as AI scribes, which include ambient listening to a medical visit and then create a written summary for provider to use for their documentation, will have access to PHI. Thus, it is important to ensure that the use of AI tools that have access to visit transcripts or chart information is thoroughly vetted and complies with regulatory laws (APA, 2025b). 

AI can be used in healthcare to aid providers in clinical decision making. The majority of this literature relates to chronic physical health conditions such as diabetes and includes non AI programs (Alnattah et al., 2025). As such, research on AI clinical decision support tools in mental health is in its infancy. A scoping review focused on AI clinical decision support found that studies do not reveal how AI has been trained to assist in clinical decision making (Auf et al., 2025). Moreover, AI tools available are not specifically developed for facilitating shared decision making between client and mental health provider (Auf et al., 2024, Auf et al. 2025). This is important because shared decision making in mental health is rooted in trust, alignment with client values, and communication, while AI models are primarily designed for triage, workflow support, screening, and prediction (Auf et al, 2024, Auf et al. 2025). Lastly, studies are lacking on the implementation of AI clinical decision support, outlining how and when mental health providers can use this tool within existing treatment workflows (Auf et al., 2025). This can leave the creation of guidelines around clinical decision support up to individual users when there should be larger discussion and policy around its use.

Concerns Related to the Client

Client confidentiality and safety are significant concerns. Issues such as how AI servers store client data, whether the servers are secure, and whether or how often client data is deleted from servers are just some examples of AI technicalities that can significantly impact client confidentiality. The lack of transparency regarding how user data is used or stored remains a major hurdle for further development of digital mental health tools (DMHI; Torous et al., 2025). Indeed, a recent scoping review found there is a meager number of studies on AI that focus on safety and privacy (Hua et al., 2025). Frighteningly, no studies were found on security, accountability, and transparency. AI tools typically use existing proprietary models, which contribute to concerns about being able to examine external validity, reliability, and safety (Hua et al., 2025).  Proprietary AI models limit the ability to assess external validity, reliability and safety, as they are largely accessed via cloud application programming interfaces (API) that communicate within a cloud computing landscape (the extent of which is outside the scope of this discussion) and infrastructures that do not allow independent auditors to inspect and/or reproduce results (Sugihara, et al, 2026)

While AI has been trained on billions of datapoints, it still struggles to navigate complex social reasoning (Torous et al., 2025). This is likely because the datasets AI models are trained on internet content such as social media, which presents a limited subsample of the general population. For example, much social media content about mental health has been found to have misinformation (Hudon et al., 2025). This can expose gen AI to harmful bias (King, 2022; Torous et al., 2025). Therefore, when AI is used by clients, it would be beneficial for psychotherapists to encourage clients to have collaborative discussions about their use of AI. Mental health providers should provide psychoeducation to clients on how to safely engage with AI around mental health, including how to critically evaluate information; what the limits of AI are as it pertains to mental health advice; and encourage discussion of AI-generated mental health information in psychotherapy (APA, 2025). This is especially important as many AI users do not talk to others about their use of AI for mental health reasons, including medical providers that they have seen in the past 6 months (McBain et al., 2026).

AI tools can potentially provide information that can be counter to psychological treatment. For example, AI has been found to violate APA ethical principles and goes against standard psychotherapy practice when communicating with users (Iftikhar et al, 2025). Researchers observed AI demonstrating a lack of contextual understanding; poor therapeutic collaboration; deceptive empathy; bias and discrimination; and a lack of adequate safety and crisis management. This occurs even when an AI model has been prompted to follow evidence-based psychological interventions (Iftikhar et al, 2025). This is especially dangerous, as some AI tools have presented themselves as being like a therapist. There have even been instances of AI purporting that they are licensed psychotherapists when they are not (Iftikhar et al, 2025). AI users can easily be deceived and believe AI tools are adequately able to provide therapeutic support or guidance, and not understand the limitations of this technology (Khawaja & Belisle-Pipon, 2023).

Impact and Outcomes of Utilizing AI

There is a history of negative outcomes when those with existing mental health concerns or persons seeking mental health support use AI (Augustin et al., 2026; APA, 2025a). AI use is especially a concern with vulnerable populations such as adolescents and those with mental illness or mental health risk concerns, as their mental health can be exacerbated by AI use. A recent study revealed that, across AI programs, AI has difficulty discerning between low-risk and high-risk suicide-related questions. AI programs also show much variability when confronted with intermediate risk questions with suicide content (McBain et al., 2025). AI is not equipped to handle safety risks.

Still, the manner in which AI generally communicates with users is concerning even when there is no safety risk. Another recent study examined AI-associated psychosis or delusions, which occur when users develop delusional ideation through the use of AI. The researchers identified how the delusions result from a culmination of AI’s use of mirrored language, hyperfocus on personalized content, and sycophancy, causing what they termed an “amplification spiral” (Augustin et al., 2026). This amplification spiral has been found to exacerbate or cause delusional ideation in people with cognitive vulnerabilities and/or people predisposed to psychosis (Augustin et al., 2026). This tendency for mirrored language, hyperfocused personalized content, and sycophancy is likely just as dangerous for clients struggling with ruminating or obsessive thoughts.

Indeed, some of the aforementioned communicative concerns observed in AI chatbot use may arise for individuals experiencing anxiety, depression, and OCD in that they can exacerbate maladaptive feedback loops that inhibit recovery and symptom reduction. As such there is additional concern about the use of AI by users with OCD or disordered thinking, as well as by adolescent users and persons who are socially isolated (APA, 2025a). Moreover, AI has been found to be lacking in a number of important psychotherapy interventions such as diagnostic accuracy, ability to tailor interventions to a client, particularly based on their past experiences, current context, and cultural identities, and ability to develop therapeutic alliance, and engage clients emotionally (Iftikhar et al, 2025; Wang et al., 2025).

Clients may use AI in tandem with psychotherapy or other psychological interventions delivered by a professional which may confuse therapeutic boundaries and alliance. AI may be more readily available for a client than formal mental health care services, which may cause clients to become more dependent on AI tools that are not as skilled as a human provider and not informed by health care laws and practice guidelines. This can detrimentally affect a client’s outcomes. AI platforms are not independently accountable to professional licensing boards, aligned with each state’s scope-of-practice requirements, informed consent standards, confidentiality and privacy regulations, mandated reporting obligations, documentation requirements, or established standards of clinical care. To worsen matters further, utilizing these tools alongside traditional psychotherapy can blur lines, leading to therapeutic misconception, which is the phenomenon in which clients underestimate AI’s restrictions and overestimate its ability to provide therapeutic support (Khawaja,et al., 2023).

Concerns Related to Psychotherapeutic Interventions

Given the fast pace at which AI has become part of society, there is naturally a variation in comfort in using AI, among both clients and mental health providers. Some may rely too much on AI and others may be fearful of it. Some may have more access to or understanding of AI than others. These differences may ultimately lead to disparities in who accesses AI tools for mental health and therefore cause disparities in client outcomes. Some therapists may be well meaning but not understand the risks involved with their clients using AI and/or the risk involved in not assessing for their AI use. The more AI is accessed by clients daily, the more clinicians need to be informed, educated, and attentive to how they utilize these tools, be it crisis guidance, clinical decision making, symptom management, or etc.  If clinicians do not engage in this type of discussion, there may be instances in which identifying misinformation, confidentiality breaches, or the need for more intensive services are missed.  Consumers can be educated about AI limitations as well as appropriate and inappropriate use.

As mentioned earlier, AI has an inherent misunderstanding of psychological concepts. The overwhelming majority of DMHIs are developed in high income countries (HIC) and thus are trained in English/European languages and cultures. Most studies have notincluded diverse samples. The potential of DMHIs to increase access to and benefit from care is limited by the fact that ethnically/racially minoritized people are underrepresented in clinical research on AI (Kodish, et al., 2022, Ramos, et al., 2022; Robinson, et al., 2024). AI trained on older data especially may be likely to misdiagnose or overpathologize BIPOC clients (Torous et al., 2025).

While there is some preliminary promise for apps that are focused on mild depression and anxiety or sub-clinical needs, there remains concerns for the use of apps for users with existing mental health symptomatology and there is little data on the efficacy of AI for various mental health concerns. (APA 2025a; Hua et al., 2025; Linardon et al., 2024). Many studies do not provide detail on their samples’ mental health status or differentiate between mental health diagnoses or conditions, further adding to confusion about the applicability of AI for mental health care. Many studies are conducted by researchers outside of mental health or medical fields. Moreover, of the studies available, data shows that AI has limited clinical effectiveness (Hua et al., 2025).

Human guidance is shown to improve the effectiveness of DMHIs, and is recommended for general AI tools as well. As such, human oversight, also known as keeping a human in the loop, is strongly recommended by the APA and researchers (APA, 2025a; Ni & Jia, 2025). Clients would benefit from evaluation by a psychologist or other mental health provider who then, depending on the client’s specific needs and preferences, could recommend DMHIs that include AI with clear guidance, guardrails, and monitoring.

Lastly, AI tools and DMHIs are still in the early stages of research, as the vast majority of studies on AI tools are pilot studies. Studies on mental health care and AI lack standardized outcome measures and adequate sample sizes (Auf et al., 2025; Hua et al., 2025). AI tools have not yet been studied enough to show their applicability to be used with the general public or for mental health treatment. Further research is warranted to understand whether AI tools are safe for inclusion in mental health care along with optimal timing and amount of use (APA, 2025; Torous et al., 2025). Typically, psychotherapies and other mental health treatments go through rigorous testing and validation before use by psychologists and other qualified professionals. AI is still in early development and yet is being used by society at large, causing many to use a tool that is not created for mental health care or not yet fully vetted to provide mental health support (McBain et al., 2026). Some argue that AI that is created for mental health support in particular should also be avoided due to these aforementioned concerns (Iftikhar et al, 2025).

Benefits of AI Use in Mental Health

Despite the numerous negatives, the use of AI technology is not without benefits. These can be grouped into benefits of accessibility, efficiency, and treatment including measurement and monitoring, and symptom reduction. Still, clients and mental health providers are urged to proceed with caution and to seek continuous education on the changing landscape of AI. When used in conjunction with professional psychological care, AI can increase clients’ access to resources and improve availability of support. For mental health providers, it can help ease the burden of administrative tasks.

Benefits Related to Accessibility

AI has the potential to expand access to care as it can eliminate some traditional barriers such as transportation difficulties, geographical distance, financial barriers, and mobility issues. It can therefore improve access within areas that are known treatment “deserts” and reduce barriers for those with disabilities. Aside from these environmental barriers, there continues to be a shortage of mental health professionals to meet societal demand, and AI tools can be helpful in mitigating this (APA, 2025). AI mental health apps have the potential to reduce wait times and increase user engagement, they can be used as part of a pre-treatment phase involving referral, assessment, and also contribute to an initial working diagnosis (Ni & Jia, 2025).

AI tools can help clients who have personal or social barriers that keep them from seeking mental health care such as anxiety or fear, stigma, shame, lack of affordable services in their area, mistrust in healthcare systems, and desire to work on their concerns independently. Indeed, some have reported that their therapeutic alliance with AI is comparable to in-person sessions (Saha, et al., 2022). However these studies thus far are limited to just some types of psychotherapy, are with limited clinical populations, and do not examine psychotherapy with human licensed psychotherapists as a comparison group. Some compare AI psychotherapy to no treatment at all (Heinz, et al., 2025).  AI has been found to have promising ability to provide psychoeducation, perform emotional awareness, and help with symptom tracking (Ni & Jia, 2025; Wang et al 2025).

With regard to symptoms, rates of anxiety, depression, psychosis, and substance use are lower among Black, Latinx/e, and Asian American adults compared to White adults (Alvarez et al., 2019; CBHSQ, 2020; Substance Abuse and Mental Health Services Administration [SAMHSA], 2015). Yet, rates of unmet mental health need are significantly higher among racially and ethnically minoritized people (Alegría et al., 2008; Breslau et al., 2005, 2006; SAMHSA, 2015). Increased access to psychological care through AI can potentially reduce this gap, although within the context of the risk of well-documented embedded racial bias. Indeed Black youth have been found to use AI for mental health needs five times more often than white counterparts (McBain et al., 2026). AI-driven chatbots or other AI tools also make mental health support more accessible and affordable as it is low cost and can be accessed by anyone with the means and technological ability.

Benefits Related to Clinician Efficiency

AI can improve task efficiency for psychologists and other mental health providers. Tools like note-taking assistants that can create session transcripts, AI scribes, and automated patient scheduling systems streamline administrative tasks, allowing therapists to reduce their administrative burden. Often, providers must either reduce one-on-one time in session (e.g., 45-or 50-minutes vs, 60 minutes) or reduce the number of therapy slots available in their schedules, if able, to accommodate time required for these administrative tasks. Employing AI tools can allow them to devote more time and focus on direct client care. Likewise, AI chart summarization tools can ease administrative burden. Such tools can provide thorough summaries of extensive medical records that may be very time consuming for a provider to review. These documentation tools can thus increase provider productivity while reducing administrative burden. Many providers find themselves working beyond their scheduled work hours to complete documentation. This can lead to a work/life imbalance and contribute to burnout. It is important to note that such AI-derived documentation must be reviewed by the clinician to ensure accuracy. For example, a human should review AI summaries of a client’s record to ensure that the information is accurate. AI tools should be able to link back to the portion of the client’s chart or session transcript where the information originated, to facilitate confidence in the accuracy and relevance of the information obtained.

AI psychotherapist guides can assist providers in their assessment and/or treatment of clients. There are numerous and valuable options that these guides can offer such as: providing treatment flowcharts or sample treatment plans, providing feedback based on treatment status, helping with diagnostic rule outs, predicting symptom or risk severity, providing suggestions regarding what areas of focus to assess next, and predicting or monitoring treatment outcomes and progress. Such AI tools can complement psychological care by ensuring that all aspects of a client’s treatment are being monitored and evaluated in a structured and comprehensive manner (Ni & Jia, 2025).

Benefits Related to Treatment

Measurement based care is important in clinical care as it provides data to inform and validate treatment approaches. AI can assist with gathering data to help monitor client needs and outcomes. These tools can help provide behavioral health screeners or measures, with data tracking, and monitor biomarkers of mental health such as blood based molecular indicators, structural and functional brain pattern analysis, and heart rate variability (HRV) (Baydili,, et. al., 2025, Lee, et al., 2021).  For example, AI can analyze HRV via enhanced ECG to identify changes in emotional regulation and stress states, as well as track inflammatory proteins and genetic expression patterns, and discriminate between bipolar and unipolar depression via brain imaging features (Baydili,, et. al., 2025, Lee, et al., 2021). AI trackers can potentially collect and analyze vast amounts of data from patient interactions, wearables, and other information sources to provide therapists with insights they might miss. This data can yield a suggested diagnosis or areas of concern, predict treatment needs or outcomes, and suggest personalized care plans (Sheikh et al., 2021, Woll et al., 2025).

AI tools can be used as a constant support between visits, potentially increasing the use of or engagement with treatment. Increasing support, including measurement based care, in this manner can assist psychotherapists with clients’ symptom reduction. Examples of how AI can be utilized include tracking psychotherapy homework and facilitating clients’ practice of skills/homework. AI chatbots and other digital mental health interventions (DMHIs) can provide interventions assigned by the provider such as cognitive-behavioral therapy (CBT) or other psychotherapy exercises, psychoeducational content, and offer clients support and tips for skill-building between sessions. DMHIs have been found to be effective for a range of mental disorders ranging from anxiety, depression, substance use (Bassilios, et al., 2022; Pineda, et al., 2023; Lattie, et al., 2022). DMHIs can be delivered in self-guided or therapist-assisted formats.

Incorporating the information obtained from AI technology can facilitate thoughtful and individualized treatment planning. AI supports can capitalize on biometric and self-reported data, allowing psychotherapists to provide care that is unique to a patient’s current mental state and concerns. This personalization, or digital phenotyping, has shown promise from numerous pilot studies focused on mental health disorders (Ni & Jia, 2025). And as mentioned throughout, digital phenotyping must also include therapist oversight.

 AI screening and/or symptom assessment can support clinical decision making and guide treatment selection. Such tools can help “fast track” the triage and treatment of clients within a health care agency (Ni & Jia, 2025). AI can also help monitor clients’ progress and safety during and after treatment. Outside of direct treatment, AI can be used as a population health tool, aiding in preventative efforts at the population level and potentially offsetting the strain on mental health care systems (Ni & Jia, 2025).

AI can serve a variety of benefits to clients as an adjunct to professionally delivered mental health care. These uses should be prescribed and monitored by a psychologist or other mental health professional. AI technology can offer real-time client feedback, psychoeducation, and assistance with home practice in between sessions. This can come in the form of adaptive skills coaching, exposure hierarchy guidance and tracking, and guided cognitive reframing. Clients can receive personalized homework practice reminders, behavioral activation planning and tracking, relapse prevention planning and tracking interventions. AI can potentially transform therapy home practice into gamified activities that serve as positive reinforcement of adaptive behaviors or coping skills (Cheng et al., 2023).. 

AI can also be advantageous in direct therapeutic intervention. Pilot studies have shown efficacy in AI use for diagnosis, risk prediction, relapse prevention, and with providing psychoeducation (Ni & Jia, 2025). However, this area of research has not extended beyond preliminary pilot studies (Torous et al., 2025). Contrarily, there are additional studies from a systematic review that revealed concerns for the use of AI in these same areas (Wang et al 2025). As such, both clients and providers should always proceed with caution, thoroughly researching and vetting AI tools prior to adopting them for personal use.

Recommendations and Conclusion

Responsible use of AI in clinical practice requires attention to a number of domains including ethics, direct clinical care, client use of AI outside of clinical care, and human oversight. When considering ethics and professional guidelines at the macro level, psychologists and other mental health providers should leverage their expertise by continued support and advocacy for the topic of AI that has been incorporated into the APA Ethics Code or encouraging the APA to develop official practice guidelines for the use of AI. Efforts can also center on promoting legislation that limits the manner in which AI is represented, offering a visible disclaimer indicating that the technology is not itself a licensed mental health clinician or qualified platform to provide professional mental health advice. Lastly, psychologists and other mental health providers can advocate for ethical AI development and AI safeguards.

Meso and micro level AI technology use should take into account how the intersection of AI and clinical practice impacts key aspects of therapeutic treatment and care. For example, steps to ensure that AI used in clinical care respects not only client confidentiality, but also state and national laws and regulations. Clear and transparent approaches to selecting and employing AI tools that are mindful of and intended to reduce bias and harm should be outlined. Prior to the use of any AI tools, it is imperative that informed consent from clients be obtained. Following this process, all clients should be equally given opportunities to use and decline AI tools within their treatment. In order to ensure this, clients should be provided education on AI tools used during assessment and treatment including the role, risks, and benefits of AI tools.  

In order to utilize AI technology as effectively and ethically as possible in direct clinical care, psychologists and other mental health providers should always acquire adequate training and understanding of the AI tools used in their clinical practice. It is crucial to stay informed on current risks and benefits of their use. Clients should be equipped with adequate information on the reasons a provider has for including AI tools in clinical practice and be informed that general purpose generative AI chatbots are not meant to replace formal mental health treatment.

Often the question is not about whether clients use AI but how they use it. AI use should be collaboratively explored and openly discussed as it relates to treatment with an emphasis on shared decision making, transparency, and continued conversation.  Psychologists and other mental health providers should regularly ask clients about their use of AI or wellness apps and encourage them to discuss any AI generated advice they have received in order to help dispel any misinformation or harmful recommendations, as well as to amplify helpful advice or information. The client should be invited to engage in discussions designed to maintain human oversight; and clinicians may give recommendations where possible if a client chooses to use AI tools outside of the provider’s treatment plan or clinical recommendation. This is especially critical given the possibility of risks to client safety and potential harmful outcomes. Such transparent and authentic conversations should be regularly implemented to support informed, collaborative decision making with regard to AI use. Clients and therapists can work together to discuss how AI is being used, have frank talks about limitations and benefits, and indicators for when additional professional support may be required.  This collaboration is particularly important for vulnerable populations for whom AI use should be thoughtfully explored within the context of the client’s treatment goals, needs, and preferences. Psychologists and other mental health providers must be sure to clarify whether clients explicitly understand AI’s limited role as an aid and that AI is not a replacement for clinical mental health treatment. Finally, clients should always be educated never to rely on AI tools for professional judgment. In the end, the goal is not to inhibit the use of AI but to encourage transparent, clinically responsible, and informed use in which the psychologist or other qualified professional holds responsibility for clinical treatment and judgment while AI remains a supplemental resource.

References: Practice

Aafjes‐van Doorn, K. (2025). Feasibility of artificial intelligence‐based measurement in psychotherapy practice: Patients’ and clinicians’ perspectives. Counselling and Psychotherapy Research, 25(1). https://doi.org/10.1002/capr.12800

Alnattah, A., Jajroudi, M., Fadafen, S. A. N., Manzari, M. N., & Eslami, S. (2025). Artificial Intelligence in Clinical Decision-Making: A Scoping Review of Rule-Based Systems and Their Applications in Medicine. Cureus, 17(8), e91333. https://doi.org/10.7759/cureus.91333

American Psychological Association. (2025, July). Ethical guidance for AI in the professional practice of health service psychology. https://www.apa.org/topics/artificial-intelligence-machine-learning/ethical-guidance-ai-professional-practice

American Psychological Association. (2025a, November). APA health advisory on the use of generative AI chatbots and wellness applications for mental health. https://www.apa.org/topics/artificial-intelligence-machine-learning/health-advisory-ai-mental-health

American Psychological Association. (2025b, December). Guidance for the evaluation of AI scribes. American Psychological Association Office of Health Care Innovation. https://www.apa.org/topics/artificial-intelligence-machine-learning/evaluating-ai-scribes

Auf, H., Svedberg, P., Nygren, J., Nair, M., & Lundgren, L. E. (2025). The Use of AI in Mental Health Services to Support Decision-Making: Scoping Review. Journal of medical Internet research, 27, e63548. https://doi.org/10.2196/63548

Auf, H., Nygren, J., Lundgren, L. E., Petersson, L., & Svedberg, P. (2025). Healthcare Professionals’ perspectives on AI-driven decision support in young adult mental health: an analysis through the lens of a shared decision-making framework. Frontiers in Digital Health, 7. https://doi.org/10.3389/fdgth.2025.1588759

Augustin, M., Pollak, T.A. & Morrin, H. Characterizing the spiral: potential mechanisms in AI-associated delusions. NPP—Digit Psychiatry Neurosci 4, 14 (2026). https://doi.org/10.1038/s44277-026-00065-0

Bassilios, B., Morgan, A., Tan, A., Ftanou, M., Krysinska, K., Le, L., … & Pirkis, J. (2022). Literature review of effectiveness of supported digital mental health interventions (DMHIs). Melbourne: The University of Melbourne.

Baydili, İ., Taşcı, B., & Taşcı, G. (2025). Artificial Intelligence in Psychiatry: A Review of Biological and Behavioral Data Analyses. Diagnostics, 15. https://doi.org/10.3390/diagnostics15040434

Bhattacharyya M, Miller V M, Bhattacharyya D, et al. (2023) High Rates of Fabricated and Inaccurate References in ChatGPT-Generated Medical Content. Cureus 15(5): e39238. doi:10.7759/cureus.39238

Bloch-Atefi, A. (2025). Balancing Ethics and Opportunities: The Role of AI in Psychotherapy and Counselling. Psychotherapy and Counselling Journal of Australia. https://doi.org/10.59158/001c.129884 

Cheng, C., & Ebrahimi, O. V. (2023). Gamification: a Novel Approach to Mental Health Promotion. Current Psychiatry Reports, 25, 577 – 586. https://doi.org/10.1007/s11920-023-01453-5

Cheng, M., Yu, S., Lee, C., Khadpe, P., Ibrahim, L., & Jurafsky, D. (2025). Sycophantic AI decreases prosocial intentions and promotes dependence.. Science, 391 6792, eaec8352 . https://doi.org/10.1126/science.aec8352

Hudon, A., Perry, K., Plate, A. S., Doucet, A., Ducharme, L., Djona, O., Testart Aguirre, C., & Evoy, G. (2025). Navigating the Maze of Social Media Disinformation on Psychiatric Illness and Charting Paths to Reliable Information for Mental Health Professionals: Observational Study of TikTok Videos. Journal of medical Internet research, 27, e64225. https://doi.org/10.2196/64225

Heinz, M. V., Bhattacharya, S., Mackin, D. M., Trudeau, B. M., Wang, Y., Banta, H. A., & Jacobson, N. C. (2025). Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI. https://ai.nejm.org/doi/full/10.1056/AIoa2400802

Hua, Y., Na, H., Li, Z., Liu, F., Fang, X., Clifton, D., & Torous, J. (2025). A scoping review of large language models for generative tasks in mental health care. NPJ Digital Medicine, 8(1), Article 230. https://doi.org/10.1038/s41746-025-01611-4

Iftikhar, Z., Xiao, A., Ransom, S., Huang, J., & Suresh, H. (2025). How LLM Counselors Violate Ethical Standards in Mental Health Practice: A Practitioner-Informed Framework. Proceedings of the AAAI ACM Conference on AI, Ethics, and Society, 8(2), 1311–1323. https://doi.org/10.1609/aies.v8i2.36632

King, M. (2022). Harmful biases in artificial intelligence. The Lancet. Psychiatry, 9(11), e48–e48. https://doi.org/10.1016/S2215-0366(22)00312-1

Khawaja, Z., & Bélisle-Pipon, J.-C. (2023). Your robot therapist is not your therapist: understanding the role of AI-powered mental health chatbots. Frontiers in Digital Health, 5, 1278186. https://doi.org/10.3389/fdgth.2023.1278186

Kodish, T., Lau, A. S., Gong-Guy, E., Congdon, E., Arnaudova, I., Schmidt, M., … & Craske, M. G. (2022). Enhancing racial/ethnic equity in college student mental health through innovative screening and treatment. Administration and Policy in Mental Health and Mental Health Services Research, 49(2), 267-282. 

Lattie, E. G., Stiles-Shields, C., & Graham, A. K. (2022). An overview of and recommendations for more accessible digital mental health services. Nature Reviews Psychology, 1(2), 87-100.

Lee, E. E., Torous, J., De Choudhury, M., Depp, C., Graham, S., Kim, H.-C., Paulus, M., Krystal, J., & Jeste, D. (2021). Artificial Intelligence for Mental Healthcare: Clinical Applications, Barriers, Facilitators, and Artificial Wisdom. Biological psychiatry. Cognitive neuroscience and neuroimaging, 6, 856 – 864. https://doi.org/10.1016/j.bpsc.2021.02.001

Linardon, J., Torous, J., Firth, J., Cuijpers, P., Messer, M., & Fuller‐Tyszkiewicz, M. (2024). Current evidence on the efficacy of mental health smartphone apps for symptoms of depression and anxiety. A meta‐analysis of 176 randomized controlled trials. World Psychiatry, 23(1), 139–149. https://doi.org/10.1002/wps.21183

Martin, D. J., Garske, J. P., & Davis, M. K. (2000). Relation of the therapeutic alliance with outcome and other variables: A meta-analytic review. Journal of Consulting and Clinical Psychology, 68(3), 438–450. https://doi.org/10.1037/0022-006X.68.3.438

McBain RK, Cantor JH, Breslau J, et al. (2026). AI Chatbot Use and Disclosure for Mental Health Among US Adolescents and Young Adults. JAMA Pediatr. 180(8):884–890. doi:10.1001/jamapediatrics.2026.2015

McBain, R. K., Cantor, J. H., Zhang, L. A., Baker, O., Zhang, F., Burnett, A., Kofner, A., Breslau, J., Stein, B. D., Mehrotra, A., & Yu, H. (2025). Evaluation of Alignment Between Large Language Models and Expert Clinicians in Suicide Risk Assessment. Psychiatric Services (Washington, D.C.), 76(11), 944–950. https://doi.org/10.1176/appi.ps.20250086

Miner, A. S., Shah, N., Bullock, K. D., Arnow, B. A., Bailenson, J., & Hancock, J. (2019). Key Considerations for Incorporating Conversational AI in Psychotherapy. Frontiers in Psychiatry, 10, 746. https://doi.org/10.3389/fpsyt.2019.00746

Ni, Y., & Jia, F. (2025). A Scoping Review of AI-Driven Digital Interventions in Mental Health Care: Mapping Applications Across Screening, Support, Monitoring, Prevention, and Clinical Education. Healthcare (Basel), 13(10), 1205. https://doi.org/10.3390/healthcare13101205

Pineda, B. S., Mejia, R., Qin, Y., Martinez, J., Delgadillo, L. G., & Muñoz, R. F. (2023). Updated taxonomy of digital mental health interventions: a conceptual framework. mHealth, 9, 28. https://doi.org/10.21037/mhealth-23-6

Ramos, G., & Chavira, D. A. (2022). Use of technology to provide mental health care for racial and ethnic minorities: evidence, promise, and challenges. Cognitive and Behavioral Practice, 29(1), 15-40. 

Robinson, A., Flom, M., Forman-Hoffman, V. L., Histon, T., Levy, M., Darcy, A., … & Montgomery, R. M. (2024). Equity in digital mental health interventions in the United States: Where to next?. Journal of Medical Internet Research, 26, e59939. 

Saha, M., Malik, T., Beatty, C., & Sinha, C. (2022). Evaluating the therapeutic alliance with a free-text CBT conversational agent (Wysa): A mixed-methods study. Frontiers in Digital Health, 4, 847991. https://doi.org/10.3389/fdgth.2022.847991

Saeidnia, H. R., Fotami, S. G. H., Lund, B. D., & Ghiasi, N. (2024). Ethical Considerations in Artificial Intelligence Interventions for Mental Health and Well-Being: Ensuring Responsible Implementation and Impact. Social Sciences. https://doi.org/10.3390/socsci13070381

Sheikh, M., Qassem, M., & Kyriacou, P. (2021). Wearable, Environmental, and Smartphone-Based Passive Sensing for Mental Health Monitoring. Frontiers in Digital Health, 3. https://doi.org/10.3389/fdgth.2021.662811

Sugihara, Y., Milosavljevic, A., Jankovskaja, S., & Falk, M. (2026). A review for navigating the trade-offs: evaluating open-source and proprietary large language models for clinical and biomedical information extraction. Frontiers in Digital Health, 8. https://doi.org/10.3389/fdgth.2026.1778786

Torous, J., Linardon, J., Goldberg, S. B., Sun, S., Bell, I., Nicholas, J., Hassan, L., Hua, Y., Milton, A., & Firth, J. (2025). The evolving field of digital mental health: current evidence and implementation issues for smartphone apps, generative artificial intelligence, and virtual reality. World Psychiatry, 24(2), 156–174. https://doi.org/10.1002/wps.21299

Wang, L., Bhanushali, T., Huang, Z., Yang, J., Badami, S., & Hightow-Weidman, L. (2025). Evaluating Generative AI in Mental Health: Systematic Review of Capabilities and Limitations. JMIR Mental Health, 12, e70014–e70014. https://doi.org/10.2196/70014

Woll, S., Birkenmaier, D., Biri, G., Nissen, R., Lutz, L., Schroth, M., Ebner-Priemer, U., & Giurgiu, M. (2025). Applying AI in the Context of the Association Between Device-Based Assessment of Physical Activity and Mental Health: Systematic Review. JMIR mHealth and uHealth, 13. https://doi.org/10.2196/59660

Artificial Intelligence Technologies in Clinical Supervision

John Gavazzi, Gerry Koocher, Melanie Wilcox                  

The integration of artificial intelligence (AI) technologies into clinical supervision represents a significant development in contemporary psychological training. This article examines five domains: supervisor agents, digitized virtual patients, case conceptualization support, clinical documentation, and longitudinal competency tracking. These applications vary considerably in maturity, ranging from those supported by early research to those barely tested in supervisory contexts.

Throughout, we consider the legal and ethical guardrails that supervisors must navigate. These include the Health Insurance Portability and Accountability Act (HIPAA), the Health Information Technology for Economic and Clinical Health (HITECH) Act, and, where applicable, the Family Educational Rights and Privacy Act (FERPA), as well as evolving guidance from the profession’s ethics code (APA, 2017) and practice guidelines (APA, 2024b, 2025a, 2025b). We examine inherent risks across all five use cases such as AI-generated fabrications (hallucinations), relational erosion, and automation bias. Two others are more prominent and consequential for supervision. The first is algorithmic bias, in which models encode and amplify existing societal inequities, threatening the cultural responsiveness supervisors are obligated to model and teach (Hofman et al., 2024). The second is deskilling, which occurs when trainees outsource the cognitive effort through which competence develops. We then turn to the competencies supervisors need for working with artificial intelligence: critical AI appraisal, prompt engineering, parallel process vigilance, data privacy and regulatory competence, and ethical integration.

The crux of every section is that supervisors remain the responsible clinical decision-makers, the human in the loop. This is not merely procedural: how supervisors engage with AI is itself a pedagogical act, as supervisees internalize not only the outputs their supervisors accept or question, but the stance with which they do so. Writing in the earliest days of cybernetics, Wiener (1948) argued that technological capability neither replaces nor diminishes the irreducibly human responsibility at the center of meaningful human processes. In supervision, that responsibility is relational, ethical, and developmental, and agnostic to algorithmic sophistication.

Foundational Concepts for AI-Assisted Supervision

Effective supervision has always required a demanding blend of skills, including delivering honest clinical feedback, supporting reflective practice, knowing when to intervene as a gatekeeper, and modeling empathy and self-care under pressure (Falender & Shafranske, 2021). As AI tools become embedded in training environments, these competencies are not replaced but extended. Supervisors must now develop enough skill with AI to guide supervisees in their ethical use, critically evaluate AI-generated clinical material, and maintain transparency about what happens to patient data (APA, 2025b). Supervision requires training distinct from clinical experience alone (Borders et al., 2023), and that training must now include critical appraisal of AI tools, which carry their own assumptions, biases, and limitations (Farmer et al., 2025).

Large Language Model Fundamentals

A working knowledge of large language models (LLMs) helps supervisors evaluate the tools their trainees may be using. Contemporary LLMs are built on transformer architectures that predict the next fragment of text, called a token, from statistical patterns in their training data (Vaswani et al., 2017). They generate coherent, contextually responsive output but possess no understanding or empathy, and output quality depends heavily on user prompts. The tool is only as clinically useful as the expertise of the person wielding it. This is especially true given hallucinations (the generation of plausible but factually incorrect or fabricated content; Linardon, et al., 2025), including invented citations and clinically reasonable but contraindicated recommendations. Mitigating this risk requires ongoing verification of outputs, which itself requires expertise; and, the remains supervisor ultimately responsible for clinical accuracy.

A second critical distinction is between general-purpose and fine-tuned models. General-purpose models (e.g., ChatGPT) are trained on broad internet datasets and may produce responses inconsistent with evidence-based practice. Fine-tuned models undergo additional domain-specific training on curated datasets, which substantially reduces, but does not eliminate, the limitations of general models. Fine-tuning also introduces a constraint of its own: a model optimized for one theoretical orientation or training level should not be assumed to transfer to another.

Navigating the Regulatory Landscape

Clinical supervisors integrating AI tools do not operate in a regulatory vacuum. HIPAA and the HITECH Act establish the foundational framework for protecting patient information and apply to any AI tool that accesses, processes, or stores protected health information (PHI). Such a vendor is functioning as a business associate with a formal agreement (Business Associate Agreement, or BAA); thus, an impermissible disclosure through an AI tool falls within HITECH’s breach notification obligations despite the rule predating the technology, and using such a tool with PHI may constitute a violation regardless of intent (Office for Civil Rights, 2013). FERPA adds further complexity in university-based settings, where supervisee records may qualify as protected educational records. The boundary between FERPA and HIPAA depends on where supervision occurs and what information is submitted to the AI tool. A cascade of boundary risks may exist. The AI user will want the assurance of a BAA and may also need to ensure that the AI vendor has similar agreements with any infrastructure or cloud computing platforms used.

APA’s Ethical Principles of Psychologists and Code of Conduct (2017) establishes that core obligations, including competence, informed consent, confidentiality, and avoidance of harm, remain in force regardless of technological novelty. The Code was nonetheless written before these technologies existed and should be treated as a necessary but not sufficient guide. The Guidelines for Clinical Supervision in Health Service Psychology (APA, 2025a) and Ethical Guidance for AI in the Professional Practice of Health Service Psychology (APA, 2025b) do address supervisory competence, data privacy, informed consent, and the obligation to maintain human decision-making authority; and, the Companion Checklist (APA, 2024c) offers a practical starting point for evaluating tools before adoption. Taken together, these frameworks provide no simple compliance roadmap. The most important skill supervisors can develop is the capacity to reason across them, recognizing where they converge, where they conflict, and where they leave space that professional judgment must fill.

AI in Clinical Supervision

The above issues recur across applications, and each application rests on the supervisor remaining the accountable human in the loop. This section examines five use cases, each at a different stage of readiness and each carrying distinct implications for the supervisory relationship.

Supervisor Agent: Fine-Tuned LLMs

Two recent studies illustrate the current state of supervisor agent technology, and the difference between them is instructive. Cioffi and colleagues (2025) presented an anonymized clinical case to three sources: a naïve ChatGPT-4 session, a ChatGPT-4 session primed with a prompt establishing Gestalt expertise, and an experienced human supervisor. Gestalt psychotherapy trainees rated the three sets of feedback blind. The primed model was rated significantly higher than both the naïve model and the human supervisor on professional orientation and adaptability, the dimension covering ethical grounding, contracting, and fit to the supervisee’s developmental level. On the relational and emotional dimension, the primed model outperformed the naïve model but was commensurate with the human supervisor. The human supervisor was rated significantly higher on constructively identifying areas for improvement, notably among the more consequential supervisory functions. Given their results, the authors propose a blended model with AI supporting self-reflection and countertransference management between sessions while human supervisors prioritize critical, growth-oriented feedback.

Two cautions apply. First, the authors attribute the model’s ratings on empathy to rhetorical facility rather than empathic understanding, consistent with the account of LLM output offered earlier. Second, authors note that their design used a single case evaluated by a small, homogeneous sample of trainees from one program. The study demonstrates feasibility, not comparability. Xu and colleagues (2025) took a different approach, moving from prompting to genuine fine-tuning. They trained models on a dataset of common therapeutic mistakes to locate problematic therapist utterances, classify error types, and generate constructive feedback, demonstrating the feasibility of scalable, immediate, and standardized feedback for novice trainees. The contrast is itself informative: prompting adjusts a general model’s behavior within a session, while fine-tuning alters the model, and the two carry different implications for reliability, portability, and the technical support a training program would need to sustain them.

These findings must be weighed against what AI cannot replicate. The supervisory working alliance is the crux of effective supervision and influences even the relationship between supervisees and their patients (Park et al., 2019). A model can produce feedback, but it cannot model the professional behaviors trainees absorb as a supervisor navigates difficult clinical moments in real time, including how uncertainty is tolerated, mistakes are acknowledged, and concern for patients is expressed under pressure. Trainees learn by watching senior professionals practice. Whatever technology is used, the human supervisor retains full professional and ethical responsibility for the clinical decisions and actions of the supervisee (APA, 2017, 2025b).

The Use of Simulated, Digitized Virtual Patients

Digitized virtual patients (VPs) offer trainees structured environments in which to practice clinical skills before working with real patients. Psychology has adopted this technology more slowly than nursing and medicine (Sanz et al., 2025), yet the evidence positions VPs as valuable pedagogical tools. They reduce ethical and practical tensions of training with real patients, adapt well to online formats, and can be customized to trainee development (Zalewski et al., 2023). A scoping review by Imam Hossain et al. (2024) found that VP simulations show promise as a supplement for building clinically relevant skills. However, the review identified three notable limitations: psychology lacks standardized guidelines for VP use, attrition averaged approximately 30%, and most simulations relied on scripted response options rather than open dialogue (the latter has been addressed by newer models). AI-enhanced VPs now support interviewing, differential diagnosis, and empathic responding with greater realism (García-Torres et al., 2025).

One under-explored dimension involves cultural and structural responsiveness. Most VP platforms have been developed within white, Western, English-language clinical frameworks, and research consistently demonstrates that racism and other forms of societal bias replicate across AI technologies (Hofmann et al., 2024). Such platforms may not recognize or respond appropriately to culturally or structurally grounded interventions. Not only might the trainee be penalized, but the simulation may disincentivize such interventions through the absence of response. The inverse risk deserves equal attention: a VP built to portray a patient from a marginalized group may render that patient through stereotype, so that trainees practice responding to a caricature and are rewarded for doing so (Bouguettaya et al., 2025). Intentional representation, and the challenging of dominant oppressive frameworks, should therefore be treated as core design requirements rather than refinements. Longitudinal research on whether VP-trained competencies transfer to real clinical settings will also be essential before these tools become standard components of training.

AI-Assisted Case Conceptualization in Supervision

Case conceptualization requires supervisees to integrate presenting symptoms, developmental history, cultural context, structural considerations, assessment data, and theoretical orientation into a coherent formulation (Wilcox, Pérez-Rojas, et al., 2024). Emerging evidence suggests LLMs may offer meaningful support, particularly for novice clinicians, by generating differential formulations and surfacing hypotheses a trainee may not have considered (Kim et al., 2025). Used this way, AI tools functions like a consultation resource that the human supervisor reviews and contextualizes.

Several concerns warrant attention. The first is over-deference. Research on automation bias shows that people uncritically accept algorithmic outputs presented in confident, fluent language (Romeo & Conti, 2025), and even practicing psychologists sometimes adopt computer-generated interpretations without deliberation (Farmer et al., 2025). Supervisees and supervisors work with layers of nuance, including nonverbal and paraverbal communication as well as behavioral observation, that cannot easily be transferred into LLM inputs. LLMs are unsuited to high-risk scenarios such as suicide or other dangerousness assessment (Balan & Gumpel, 2025).

Second, general-purpose LLMs are trained predominantly on white, Western, English-language datasets that systematically underweight cultural identity, social location, and systemic factors such as racism and poverty (Strand & Osanami Törngren, 2025), and they may reify professional power structures and diagnostic disparities (Bouguettaya et al., 2025). Fine-tuning and skilled prompting may partially help (Tao et al., 2024; Xie et al., 2024), but cultural and structural responsiveness should be treated as required evaluation criteria rather than desirable features.

Third, case conceptualization is not only a product but also a developmentally significant process and routinely outsourcing that integration is where the deskilling risk becomes concrete (Choudhury & Chaudhry, 2024). Supervisors would do best to use AI-generated formulations as objects of critical analysis rather than as starting templates. Absent published outcome studies or standardized guidelines, appropriateness should follow a trainee-by-trainee assessment model.

AI-Assisted Clinical Documentation

AI documentation tools use ambient listening and natural language processing to transcribe and summarize sessions into common structured formats. Adoption remains limited, with only a minority of physicians currently using them (Rotenstein et al., 2026); but, the potential to reduce administrative burden is significant. Performance is uneven. Some tools struggle with accents, multiple speakers, or complex presentations, and comparable data from mental health settings are not yet available (Alpert et al., 2025). Successful implementations have used phased rollouts, live training, and, critically, mandatory human review and editing of all AI-generated notes (McCrudden et al., 2026). Research also suggests AI scribes may overemphasize symptoms while under-documenting interventions (Castro et al., 2026),  which could result in  trainees or clinicians reading back their AI-drafted notes and absorbing a symptom-focused account. Supervisors should review notes specifically for treatment planning and intervention rationales.

Ambient listening also raises relational questions (Haber et al., 2024). Trainees may orient their language toward documentable content, and patients may shift what they disclose. Supervisors should help trainees maintain genuine therapeutic presence and prepare them to introduce the technology to patients by explaining in plain language what it captures, how recordings and transcripts are stored, and that patients may decline its use without consequence. Note-writing is itself a cognitive and reflective exercise through which durable clinical reasoning is consolidated; thus, programs may consider delaying AI scribe use until trainees demonstrate baseline documentation competency.

Supervisors must treat AI-generated notes as requiring active cultural and antioppressive critique, not merely factual verification (Falender et al., 2014). Consent for AI tools is a baseline obligation in any clinical setting (APA, 2025b). In training, supervisors must additionally ensure that supervisees can hold that conversation competently rather than treating it as paperwork.

Assessing the Trainee: Competency Tracking

Traditional trainee evaluation relies heavily on end-of-semester ratings and supervisor recall—all vulnerable to recency effects, halo effects, and leniency bias. Improvement efforts, such as the competency tracking tool validated by Lawrence et al. (2024), aim to address these limitations. AI-driven analytics can assist by automating and enriching data collection. VP interactions can generate behavioral data from a trainee’s first simulation (Morrison et al., 2025), and as trainees advance, ambient scribes and transcript analysis can track observable competency markers across placements, such as talk time distribution and explicit negotiation of goals and tasks (Brown et al., 2026). Alliance quality itself remains outside what these tools can assess. Natural language processing can quantify concrete indicators of interpersonal skill, including reflection-to-question ratios and talk-time percentages, providing data on patient-centered dialogue (Orrù & Mannarini, 2026).

Nonetheless, evidence on AI-generated educational feedback remains mixed (Mpolomoka, 2025), and inaccuracies and algorithmic bias can distort competency ratings (Landers & Behrend, 2023). Trainees from underrepresented backgrounds already encounter racism and other forms of oppression in training (Wilcox, Farra, et al., 2024), and an evaluation system that encodes bias compounds an existing burden. AI-generated competency ratings must therefore be reviewed critically and should never serve as the sole basis for high-stakes decisions such as remediation or advancement. The supervisor remains the final interpreter of the data, and the supervisory relationship, not AI, remains the center of the feedback process.

Supervisor Competencies

Each application carries genuine promise and equally genuine risk. Just as evidence-based practice requires supervisors to develop new skills alongside existing ones, AI integration also demands an expansion of supervisors’ repertoires. Foundational supervisory skills supported by decades of research remain paramount, but five additional competencies are notable.

Critical AI Appraisal (Beyond Basic Literacy)

Understanding how AI systems function is necessary but insufficient. This calls for an evaluative framework, not technical understanding alone — a higher-order capacity to evaluate when, why, and for whom an AI-generated recommendation can be trusted.                                                  Supervisors must have the ability to evaluate AI outputs against evidence-based practice. LLMs may generate plausible-sounding but contraindicated suggestions; thus, supervisors need to identify embedded oppression (e.g., racism, sexism, classism), and diagnostic categorization (Hofmann et al., 2024), drawing explicitly on multicultural, structural, and social justice competency frameworks (Wilcox, Pérez-Rojas, et al., 2024), positioning critical appraisal as an extension of cultural humility. Appraisal also extends to patients’ own increasing use of general-purpose LLMs for mental health purposes. Trainees should learn to ask directly about such use, much as they would inquire about risky practices, which opens the door to psychoeducation about AI limitations. Finally, supervisors must recognize the boundaries of AI applicability, since tools validated on majority-population datasets may perform poorly and cause harm with culturally complex or marginalized populations (Bouguettaya et al., 2025).

Prompting as a Supervisor Competency

Prompt engineering, the design and refinement of structured inputs to guide LLM output toward clinically useful ends, now warrants inclusion among the discrete skills developed in competency-based supervision (Falender & Shafranske, 2021). Small changes in wording can substantially affect output quality (Liu et al., 2025). Three dimensions matter. First, supervisors should teach supervisees to craft specific, context-rich prompts, as vague prompts yield generic responses lacking clinical relevance, cultural/structural responsiveness, and theory-drivenness. Second, supervisees must learn to refine prompts iteratively when outputs prove inadequate, evaluating why a prompt failed and revising accordingly—a process that mirrors the reflective practice foundational to trainee development (Falender & Shafranske, 2021). Constructing, comparing, and contrasting prompts together in supervision makes this concrete. Third, and most consequential, is supervisor modeling: sharing one’s own prompts, explaining their reasoning, and demonstrating how to evaluate results. Parallel process applies: supervisors using AI privately miss a teaching opportunity and may inadvertently model nondisclosure, increasing the likelihood that supervisees will use AI without telling them. Making expert reasoning visible is a well-supported mechanism for developing clinical reasoning in novices.

Parallel Process Vigilance

Parallel process, the replication of relational patterns from therapy within supervision (Tracey et al., 2011) may plausibly extend to the supervisee’s relationship with AI. AI increasingly functions as a third presence in the clinical training environment (Haber et al., 2024), a dynamic that warrants empirical investigation, though early implementation reports hint at its relevance (Sheperis & Sadeh-Sharvit, 2023). A supervisee may treat AI output as a reliable, even superior, clinical voice. If the parallel holds, this deference may reproduce a dynamic already present elsewhere in the system, one in which the patient defers to the supervisee and the supervisor defers, often without examination, to technological efficiency. Automation bias research documents such over-reliance in clinical contexts. Supervisors should intervene through Socratic questioning—asking what the AI output assumed, what it omitted, and how it aligns with the supervisee’s own relational knowledge of the patient (Falender & Shafranske, 2021). Conversely, supervisors should also remain vigilant about reflexive and categorical dismissal of AI, itself a form of avoidance worth exploring. Supervisor modeling is especially consequential (Watkins, 2014); supervisor who demonstrates a balanced, critical, metacognitively transparent stance on AI and thinks aloud about how they evaluate output and when to set it aside, teaches professional judgment in vivo (Abdulnour et al., 2025).

Data Privacy and Regulatory Competence

Beyond general familiarity with HIPAA, HITECH, and FERPA, AI tools require an applied layer of regulatory knowledge, and APA’s Companion Checklist (2024c) offers a starting point. Any tool that accesses, processes, or stores PHI requires a signed BAA with the vendor before clinical use, including ambient listening tools and LLMs used with session content (APA, 2024a). Vendors’ general terms of service do not satisfy this requirement, and many widely used platforms explicitly disclaim HIPAA compliance. A related risk concerns what constitutes PHI. A prompt including a patient’s age, diagnosis, location, and enough contextual detail to make the person reasonably identifiable may qualify as PHI even without a name (Office for Civil Rights, 2025). Supervisors should teach supervisees to replace identifying details with neutral descriptors, avoid uploading recordings or transcripts to noncompliant platforms, and treat AI outputs containing patient material with the same confidentiality obligations as the clinical record. Under HITECH, impermissible disclosures through AI tools trigger breach notification obligations, and supervisees must understand their duty to report potential breaches (Office for Civil Rights, 2013). In university-based programs, AI-generated supervisee performance data may become FERPA-protected educational records, warranting consultation with institutional legal counsel. Finally, federal frameworks provide a floor, not a ceiling, since state and institutional policies may add requirements. This is especially important in interjurisdictional telehealth and telesupervision.

Ethical Integration and Boundary Management

APA has affirmed that AI should augment rather than replace human decision-making, with no reduction in psychologists’ accountability (APA, 2025b). Supervisors therefore need clear, defensible frameworks for determining which AI tools suit which supervision tasks. Documentation assistance may qualify as appropriate in many contexts, whereas AI-generated diagnostic formulations and high-risk clinical situations warrant significantly more caution (APA, 2025b). To the four bioethical principles of autonomy, beneficence, nonmaleficence, and justice (Beauchamp & Childress, 2025), the psychology adds fidelity, without which the work is not possible (Knapp & Fingerhut, 2024). The AI ethics literature adds transparency (Pillay, 2025). AI tools may serve as consultative resources but not co-supervisors. When AI-generated content supersedes the relational emphasis of supervision, supervisors risk eroding the most empirically supported mechanism of supervision, with consequences for supervisee wellbeing and professional development as well as patient care (Watkins, 2014).

Boundary management extends beyond the supervisor’s own use of these tools to what is delegated, disclosed, and taught. Supervisors should decide deliberately what work is and is not appropriate for AI support, and explain that reasoning when needed. The supervisory relationship calls for the same transparency. Supervisors who disclose their own AI use, and invite supervisees to do the same, model the openness that keeps AI a visible part of the work rather than a private workaround. Supervisees need explicit guidance on the ethical boundaries of AI use in clinical work, such as never entering identifiable patient information into noncompliant tools; disclosing AI-assisted documentation to patients where appropriate; and resisting the temptation to consult AI in place of the supervisor during moments of clinical uncertainty. Addressing these boundaries directly in the supervision agreement, rather than leaving them implicit, gives supervisees a clear standard to work from and gives supervisors a basis for intervention if that standard is not met.

Integrating the Supervisory Roles with AI Technologies

AI tools are entering clinical training environments faster than the field can evaluate them systematically. The question is not whether supervisors will encounter these tools, but whether they will engage them deliberately and critically, in a manner consistent with the profession’s deepest ethical commitments.

Across every domain examined, one principle holds without exception: supervisors remain the accountable clinical decision-makers. When a supervisor agent generates feedback, a VP produces a clinical scenario, or an AI scribe drafts a note, the supervisor’s professional judgment remains requisite. Outputs require evaluation and edits by an expert with the clinical knowledge, ethical grounding, and relational attunement no technology possesses. Wiener’s (1948) insight holds: technological capability neither replaces nor diminishes irreducible human responsibility. Decades of research identify the supervisory alliance as the most empirically supported mechanism of trainee development. That relational core is not a supplement to the technical work. It is the work itself.

Two hazards warrant direct attention. Algorithmic bias shapes every AI output encountered; and, deskilling accumulates quietly, one outsourced formulation or note at a time. Neither will resolve with better technology, and each independently necessitates the human-in-the-loop framework at this article’s center. Parallel process is not a hazard but a condition of supervision. How supervisors and supervisees engage AI may shape how the supervisee engages the patient, emphasizing the consequential nature of AI competencies. Cultural and structural responsiveness belong in a different category still. They are not hazards to be managed but the lenses through which adoption decisions must be made; such considerations cannot be ancillary.

The empirical foundation is still taking shape, and the regulatory landscape, including HIPAA, HITECH, FERPA, and the Ethics Code, predates these technologies. Where research is thin and frameworks leave gaps, the principles of autonomy, beneficence, nonmaleficence, justice, fidelity, and transparency offer a stable basis for decisions. Supervisors therefore need not wait for definitive standards to act responsibly. They can establish a written AI use policy, verify BAAs and meaningful patient consent, teach critical AI appraisal and prompt engineering explicitly, and treat AI-generated competency data as supplementary, never standalone. 

References: Supervision

Abdulnour, R. E., Gin, B., & Boscardin, C. K. (2025). Educational strategies for clinical supervision of artificial intelligence use. New England Journal of Medicine, 393(8), 786–797. https://doi.org/10.1056/nejmra2503232

Alpert, J. M., Saper, R., Boose, E., Ruff, J., Hopkins, K., Gaskins, D., Guo, N., Schneider, J. P., Giffi Scibona, A., Gutierrez, J., & Rothberg, M. B. (2025). Evaluating an artificial intelligence scribe for clinical documentation. Digital Health, 11, 20552076251395588. https://doi.org/10.1177/20552076251395588

American Psychological Association. (2017). Ethical principles of psychologists and code of conduct. https://www.apa.org/ethics/code

American Psychological Association. (2024a, August). APA guidelines for the practice of telepsychology. https://www.apa.org/practice/guidelines/telepsychology-revision.pdf

American Psychological Association. (2024b, August). Artificial intelligence and the field of psychology [Policy statement]. https://www.apa.org/about/policy/statement-artificial-intelligence.pdf

American Psychological Association. (2024c). Companion checklist: Evaluation of an AI-enabled clinical or administrative tool. APA Services. https://www.apaservices.org/practice/business/technology/tech-101/evaluating-artificial-intelligence-tool-checklist.pdf

American Psychological Association. (2025a). APA guidelines for clinical supervision in health service psychology. https://www.apa.org/about/policy/guidelines-supervision.pdf

American Psychological Association. (2025b). Ethical guidance for AI in the professional practice of health service psychology. https://www.apa.org/topics/artificial-intelligence-machine-learning/ethical-guidance-professional-practice.pdf

Balan, R., & Gumpel, T. P. (2025). ChatGPT clinical use in mental health care: Scoping review of empirical evidence. JMIR Mental Health, 12, e81204. https://doi.org/10.2196/81204

Beauchamp, T. L., & Childress, J. F. (2025). Principles of biomedical ethics (9th ed.). Oxford University Press.

Borders, L. D., Dianna, J. A., & McKibben, W. B. (2023). Clinical supervisor training: A ten-year scoping review across counseling, psychology, and social work. The Clinical Supervisor, 42(1), 164–212. https://doi.org/10.1080/07325223.2023.2188624

Bouguettaya, A., Stuart, E. M., & Aboujaoude, E. (2025). Racial bias in AI-mediated psychiatric diagnosis and treatment: A qualitative comparison of four large language models. npj Digital Medicine, 8(1), 332. https://doi.org/10.1038/s41746-025-01746-4

Brown, C., Pires, L., & Benskin, T. L. (2026). Machine-learning prediction of therapeutic alliance quality from client readiness to change and affective dysregulation. Journal of Assessment and Research in Applied Counseling, 8(1), 1–10. https://doi.org/10.61838/kman.jarac.5183

Castro, V. M., McCoy, T. H., Verhaak, P., Ramachandiran, A., & Perlis, R. H. (2026). Psychiatric documentation and management in primary care with artificial intelligence scribe use. JAMA Psychiatry, 83(3), 281. https://doi.org/10.1001/jamapsychiatry.2025.4303

Choudhury, A., & Chaudhry, Z. (2024). Large language models and user trust: Consequence of self-referential learning loop and the deskilling of health care professionals. Journal of Medical Internet Research, 26, e56764. https://doi.org/10.2196/56764

Cioffi, V., Ragozzino, O., Mosca, L. L., Moretto, E., Tortora, E., Acocella, A., Montanari, C., Ferrara, A., Crispino, S., Gigante, E., Lommatzsch, A., Pizzimenti, M., Temporin, E., Barlacchi, V., Billi, C., Salonia, G., & Sperandeo, R. (2025). Can AI technologies support clinical supervision? Assessing the potential of ChatGPT. Informatics, 12(1), 29. https://doi.org/10.3390/informatics12010029

Falender, C. A., & Shafranske, E. P. (2021). Clinical supervision: A competency-based approach (2nd ed.). American Psychological Association. https://doi.org/10.1037/0000243-000

Falender, C. A., Shafranske, E. P., & Falicov, C. J. (Eds.). (2014). Multiculturalism and diversity in clinical supervision: A competency-based approach. American Psychological Association. https://doi.org/10.1037/14370-000

Farmer, R. L., Lockwood, A. B., Goforth, A., & Thomas, C. (2025). Artificial intelligence in practice: Opportunities, challenges, and ethical considerations. Professional Psychology: Research and Practice, 56(1), 19–32. https://doi.org/10.1037/pro0000595

García-Torres, D., Fernández, C., Mira, J. J., Morales, A., & Vicente, M. A. (2025). Using AI-based virtual simulated patients for training in psychopathological interviewing: Cross-sectional observational study. JMIR Medical Education, 11, e78857. https://doi.org/10.2196/78857

Haber, Y., Levkovich, I., Hadar-Shoval, D., & Elyoseph, Z. (2024). The artificial third: A broad view of the effects of introducing generative artificial intelligence on psychotherapy. JMIR Mental Health, 11, e54781. https://doi.org/10.2196/54781

Hofmann, V., Kalluri, P. R., Jurafsky, D., & King, S. (2024). AI generates covertly racist decisions about people based on their dialect. Nature, 633(8028), 147–154. https://doi.org/10.1038/s41586-024-07856-5

Imam Hossain, S., Kelson, J., & Morrison, B. (2024). The use of virtual patient simulations in psychology: A scoping review. Australasian Journal of Educational Technology, 40(6), 76–91. https://doi.org/10.14742/ajet.9559

Kim, N., Lee, J., Park, S. H., On, Y., Lee, J., Keum, M., Oh, S., Song, Y., Lee, J., Won, G. H., Shin, J. S., Lho, S. K., Hwang, Y. J., & Kim, T. S. (2025). GPT-4 generated psychological reports in psychodynamic perspective: A pilot study on quality, risk of hallucination and client satisfaction. Frontiers in Psychiatry, 16, 1473614. https://doi.org/10.3389/fpsyt.2025.1473614

Knapp, S. J., & Fingerhut, R. (2024). Practical ethics for psychologists: A positive approach (4th ed.). American Psychological Association. https://doi.org/10.1037/0000375-000

Landers, R. N., & Behrend, T. S. (2023). Auditing the AI auditors: A framework for evaluating fairness and bias in high stakes AI predictive models. American Psychologist, 78(1), 36–49. https://doi.org/10.1037/amp0000972

Lawrence, K. A., Taylor, K., Cairns, R., Gersh, H., & McKay, A. (2024). Tracking trainee development: Preliminary validation of a tool designed to evaluate clinical psychology competencies over time. Australian Psychologist, 59(4), 329–340. https://doi.org/10.1080/00050067.2024.2353024

Linardon, J., Jarman, H. K., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models: Experimental study. JMIR Mental Health, 12, e80371. https://doi.org/10.2196/80371

Liu, J., Liu, F., Wang, C., & Liu, S. (2025). Prompt engineering in clinical practice: Tutorial for clinicians. Journal of Medical Internet Research, 27, e72644. https://doi.org/10.2196/72644

McCrudden, K. E., Swirbul, M. S., Peake, E. E., Rodio, M. J., & Padmanabhan, A. (2026). AI-powered documentation for mental health providers: Retrospective observational mixed methods study. JMIR Formative Research, 10, e84628. https://doi.org/10.2196/84628

Morrison, B. W., Kelson, J., Morrison, N. M. V., & Bennett, G. (2025). You’re virtually a psychologist: Enhancing professional psychology education through virtual client simulations. Australian Psychologist, 61(1), 22–29. https://doi.org/10.1080/00050067.2025.2547800

Mpolomoka, D. L. (2025). Utilizing artificial intelligence for assessment in higher education. Pedagogical Research, 10(3), em0243. https://doi.org/10.29333/pr/16677

Office for Civil Rights. (2013). Breach notification rule. U.S. Department of Health and Human Services. https://www.hhs.gov/hipaa/for-professionals/breach-notification/index.html

Office for Civil Rights. (2025). Methods for de-identification of PHI. U.S. Department of Health and Human Services. https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification/index.html

Orrù, L., & Mannarini, S. (2026). The role of artificial intelligence in clinical psychology: How AI and NLP systems are reshaping psychological interventions. A systematic review. Clinical Psychology & Psychotherapy, 33(2), e70242. https://doi.org/10.1002/cpp.70242

Park, E. H., Ha, G., Lee, S., Lee, S., Lee, Y. Y., Lee, S. M., & Lee, S. M. (2019). Relationship between the supervisory working alliance and outcomes: A meta-analysis. Journal of Counseling & Development, 97(4), 437–446. https://doi.org/10.1002/jcad.12292

Pillay, Y. (2025). Ethical decision-making guidelines for mental health clinicians in the artificial intelligence (AI) era. Healthcare, 13(23), 3057. https://doi.org/10.3390/healthcare13233057

Romeo, G., & Conti, D. (2025). Exploring automation bias in human-AI collaboration: A review and implications for explainable AI. AI & Society, 41(1), 259–278. https://doi.org/10.1007/s00146-025-02422-7

Rotenstein, L. S., Holmgren, A. J., Thombley, R., Sriram, A., Dbouk, R. H., Jost, M., Aizenberg, D., MacDonald, S., Kanaparthy, N., Williams, B., Hsiao, A., Schwamm, L., Murray, S., Byron, M., You, J. G., Centi, A. J., Iannaccone, C., Frits, M., Landman, A. B., . . . Mishuris, R. G. (2026). Changes in clinician time expenditure and visit quantity with adoption of artificial intelligence-powered scribes. JAMA, 335(16), 1408. https://doi.org/10.1001/jama.2026.2253

Sanz, A., Tapia, J. L., García-Carpintero, E., Rocabado, J. F., & Pedrajas, L. M. (2025). ChatGPT simulated patient: Use in clinical training in psychology. Psicothema, 37(3), 23–32. https://doi.org/10.70478/psicothema.2025.37.21

Sheperis, D., & Sadeh-Sharvit, S. (2023). Using AI-supported supervision in a university telemental health training clinic. Journal of Technology in Counselor Education and Supervision, 4(1). https://doi.org/10.61888/2692-4129.1093

Strand, M., & Osanami Törngren, S. (2025). Multiraciality and mental health: The Cultural Formulation Interview as an instrument for exploring in-between identities and third spaces. Frontiers in Psychiatry, 16, 1690109. https://doi.org/10.3389/fpsyt.2025.1690109

Tao, Y., Viberg, O., Baker, R. S., & Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models. PNAS Nexus, 3(9), pgae346. https://doi.org/10.1093/pnasnexus/pgae346

Tracey, T. J. G., Bludworth, J., & Glidden-Tracey, C. E. (2011). Are there parallel processes in psychotherapy supervision? An empirical examination. Psychotherapy, 49(3), 330–343. https://doi.org/10.1037/a0026246

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems 30 (pp. 5998–6010). Curran Associates.

Watkins, C. E., Jr. (2014). The supervisory alliance: A half century of theory, practice, and research in critical perspective. American Journal of Psychotherapy, 68(1), 19–55. https://doi.org/10.1176/appi.psychotherapy.2014.68.1.19

Wiener, N. (1948). Cybernetics: Or control and communication in the animal and the machine. MIT Press.

Wilcox, M. M., Farra, A., Winkeljohn Black, S., Pollard, E., Drinane, J. M., Tao, K. W., DeBlaere, C., Hook, J. N., Davis, D. E., Watkins, C. E., Jr., & Owen, J. (2024). Cultural humility and racial microaggressions in cross-racial clinical supervision: A moderated mediation model. Journal of Counseling Psychology, 71(4), 304–314. https://doi.org/10.1037/cou0000732

Wilcox, M. M., Pérez-Rojas, A. E., Marks, L. R., Reynolds, A. L., Suh, H. N., Flores, L. Y., McCubbin, L. D., Wilkins-Yel, K. G., & Miller, M. J. (2024). Structural competencies: Re-grounding counseling psychology in antiracist and decolonial praxis. The Counseling Psychologist, 52(4), 650–691. https://doi.org/10.1177/00110000241231029

Xie, S. J., Zhai, S., Liang, Y., Li, J., Fan, X., Cohen, T., & Yuwen, W. (2024). Cultural prompting improves the empathy and cultural responsiveness of GPT-generated therapy responses. AMIA … Annual Symposium proceedings. AMIA Symposium, 2024, 1384–1393.

Xu, C., Lv, Z., Lan, T., Wang, X., Ji, L., Cui, L., Yang, M., Shen, J., Dong, Q., Liu, X., Wang, J., & Hu, B. (2025). LLM-as-a-supervisor: Mistaken therapeutic behaviors trigger targeted supervisory feedback [Preprint]. arXiv. https://doi.org/10.48550/arxiv.2508.09042

Zalewski, B., Guziak, M., & Walkiewicz, M. (2023). Developing simulated and virtual patients in psychological assessment: Method, insights and recommendations. Perspectives on Medical Education, 12(1), 455–461. https://doi.org/10.5334/pme.49

AI in Clinical Psychology Education and Training

Firouz Ardalan and James Lichtenberg

Artificial intelligence is reshaping how training programs teach, how students’ study, and how psychotherapists are trained. AI is not new to clinical settings; what has changed is its capability (Xian et al., 2024). The emergence of generative AI (Gen AI) marks a categorically different phase. Large Language Models (LLMs) are now capable of producing novel outputs, synthesizing large amounts of information, mimicking natural language, tailoring responses to specific prompts and contexts, generating high-quality content at extraordinary speed, and working seamlessly across text, image, audio, and video modalities. This combination of capabilities makes Gen AI among the most powerful and versatile technologies available today, with serious implications for how clinicians are trained.

This segment of the SAP AI and Psychotherapy Report aims to orient readers broadly to how artificial intelligence (AI) is being used in psychology education and training: what tools are being developed, what early evidence suggests about potential benefits, concerns, and what recommendations follow for those working or learning with AI in educational and clinical settings. Companies and products in this document are named for illustrative purposes only; mention does not imply endorsement by the authors or by APA Division 29.

Overview of AI in Clinical Education and Training

One of our greatest challenges at the moment is that the adoption of AI is outpacing the research needed to evaluate its risks and benefits. A 2025 scoping review identified only 26 peer-reviewed papers on AI in psychiatric education and training, with roughly 60% published between 2023 and 2024 alone (Weightman et al., 2025). A parallel review of applications in psychology and psychiatry education screened more than 6,000 records and found only 10 that met inclusion criteria (Prégent et al., 2025). Inclusion criteria included a focus on psychiatry or psychology education, the involvement of an AI tool, model or approach, exploration of facilitators or barriers using AI, and were in English or French.  While psychology and psychiatry disciplines have long been guided by evidence-based practice, training practices have often been driven by a practical experience-based approach.  Nonetheless, the scarcity of research leaves educators and trainers needing to make consequential decisions without the benefit of accumulated evidence, or the luxury of waiting for it to accrue. Educators and trainers must nonetheless engage with what is already happening.

AI Applications to Clinical Education and Training

The AI applications already in development are broader in scope than most may realize, and their potential seems nearly limitless. At the time of writing, AI is being applied to (a) virtual patient simulations with a broad scope of abilities, (b) case-based learning, (c) case conceptualization, (d) psychological assessment, (e) student assessments, (f) mental health monitoring, (g) supervision and professional development, (h) the creation of educational resources, instructional content, and study aids, (i) clinical documentation, (j) research assistance, (k) content synthesis, and (l) policy and program development (Lee et al., 2025; Weightman et al., 2025; Prégent et al., 2025). This list illustrates the potential range and scope of AI’s influence on education and clinical training; however, it does not claim to be exhaustive. Advancements are proceeding rapidly, and new technologies are likely to have emerged by the time of publication. 

Benefits

Some of the early evidence supporting AI-enhanced clinical training is encouraging, particularly given the field’s nascent state. A meta-analysis by Kilincel et al. (2026) synthesized findings from 10 studies involving medical students, psychiatry residents, nursing students, clinicians, and psychology trainees and found consistent gains in how trainees conducted psychiatric interviews, absorbed clinical knowledge, and felt about their own competence. Across the included studies, trainees who practiced with these tools became more capable at organizing an interview, recognizing symptoms, and diagnostic reasoning. Improvements in confidence and communication skills were notably consistent across populations.

Virtual Patient Simulation 

Perhaps unsurprisingly, one of the most active areas of Gen AI development is LLM-based virtual patients (VPs). VP simulation itself predates the current generation of AI; the first known rule-based simulation was produced in 1966 (Weizenbaum, 1966). While vastly different than today’s AI, it signifies just how long technology and psychology have been intertwined. However, current AI systems can now simulate patients presenting with specific personalities, histories, symptoms, and behavioral characteristics far more convincingly than ever before. They are more conversational, adaptable, and able to respond to trainee input and adjust in real time (Kilincel et al., 2026). They are being used to train students on psychiatric interviewing, crisis intervention, suicide risk assessment, diagnostic reasoning, differential diagnosis, and exposure to rare and high-risk presentations that trainees are unlikely to encounter on placement (Kilincel et al., 2026; Weightman et al., 2025; García-Torres et al., 2025). AI systems are now capable of generating the kind of dynamic, contextually responsive conversation that clinical training has always depended on real patients to provide (Weightman et al., 2025). In addition, AI bots designed to help train students in specific modalities and intervention techniques, such as cognitive-behavioral therapy or acceptance and commitment therapy, are becoming more prevalent (Prégent et al., 2025).

Results demonstrating the benefits of using AI simulation show mixed outcomes. In a cross-sectional study involving 156 psychology students who used an AI-simulated patient to strengthen their diagnostic skills, García-Torres and colleagues (2025) found that 94% self-reported meaningful improvement in their ability to identify patients’ symptoms. Students participating in the study also reported a broadly positive experience working with the simulated patient, providing an early signal of adoptability. It’s important to note that these results are all self-reported, indicating subjective perceptions of improvement and but that their testing scores did not show any significant changes from previous cohorts who were not trained using VSP.  Sarli et al. (2024) also showed the feasibility of strengthening trainees’ empathy skills using virtual patients.

Crisis Intervention Training 

AI simulation may offer its most distinctive and clinically urgent contribution in crisis intervention training. High-acuity presentations, such as acute suicidality or psychotic decompensation, represent precisely the encounters that trainees need and may find hardest to access—and those which supervisors may actively prevent trainees from having exposure to as they balance clinical care with training. Simulation-based approaches have shown potential to improve trainee skills in crisis intervention (Richard et al., 2023). After practicing with the SEE THE PAIN simulator, clinicians in Elyoseph et al.’s (2026) study self-rated themselves as both more confident with and more willing to take on patients at risk for suicide. While self-report is a meaningful indicator, it does not indicate skill gain. The study observed the largest gains among early-career practitioners, suggesting that AI-based training may be most beneficial for those in training. Expert raters in Wang et al.’s (2024) study judged PATIENT-Ψ’s distorted thinking patterns, conversational tone, and emotional expression to be authentically patient-like. In addition, trainees self-reported that practicing with PATIENT-Ψ-TRAINER built their skill and confidence more so than textbooks, videos, or peer role-play exercises. However, as noted by the authors, further longitudinal studies will need to be conducted to assess actual skill acquisition. 

A significant benefit is the potential to ensure that every trainee encounters a full range of clinical presentations that competency-based training requires. As currently structured, a trainee’s exposure is largely determined by circumstance. The presentations to which a trainee is exposed during training depend largely on where their training site is located, who is seeking services at that time, which appointments are kept, and which presentations their supervisors prioritize. There is presently no easy solution to the problem of trainees who too often complete their doctoral training under-exposed to a diverse range of patients, leaving them feeling inadequately prepared (Wang et al., 2024). AI simulation is a genuine response to this problem, and the benefit appears to be greatest for those in training or early career. AI’s ability to be used by entire cohorts across multiple languages and cultures, and asynchronously, removes issues of access, making training more equitable than ever before.

AI simulation offers a training advantage that no other training modality can replicate at scale. For example, the creators of the SEE THE PAIN simulator built it specifically to allow trainees to engage in repeated, consequence-free practice with a patient at acute risk of suicide while receiving feedback on their interactions (Elyoseph et al., 2026). While not Gen AI, earlier work by O’Brien and colleagues (2019) demonstrated significant improvements in trainee suicide risk assessment skills among those who trained with virtual patients, further illustrating that using online virtual patients can improve skill acquisition. Advancements in Gen AI’s capabilities show promise for continued expansion and implementation of these methods.

Suicide Risk Assessment 

Several platforms illustrate the range of VPs’ capability. SEE THE PAIN (Elyoseph et al., 2026) is a Gen AI-based simulator designed to train clinicians in suicide risk assessment and build confidence working with high-risk patients. It operates across multiple languages, though its cultural and clinical validation across all of them is still limited; even so, it points to AI’s potential to standardize training across linguistic and cultural contexts in time. Wang and colleagues (2024) developed a realistic, interactive simulation (PATIENT-Ψ-TRAINER) for cognitive-behavioral therapy (CBT) training, in part, to address the gap between clinicians’ currently existing training and the complexities of acquiring real in vivo patient experience. In addition, the degree of customization these platforms achieve is notable: in PATIENT-Ψ, the conversational style of the bot can be calibrated for a range of patient presentations. Such customization empowers training programs to match a bot’s complexity to a trainee’s developmental stage and gradually increase the challenge as competency develops (Wang et al., 2024).

Case Conceptualizations, Case Presentations, Case Notes, and Documentation  

AI has a remarkable ability to synthesize information, from which students in particular can benefit. Using LLMs to facilitate editing, organizing, and drafting academic and clinical assignments can strengthen the quality of work produced. Such skills can be applied to case conceptualizations (Hsieh et al., 2024), case presentations, notes, general documentation, research methods, and tasks like grant writing. Hsieh et al. (2024) found that doctoral-level clinicians rated AI-generated case conceptualizations to be acceptable after assessing them for accuracy, completeness, and consistency. Obviously, the integrity of the work must be maintained. But when used thoughtfully and as a genuinely supportive, iterative process, it can serve as a valuable tool (Prégent et al., 2025). LLMs can also serve as study aids, producing study guides, note cards, and even apps designed to help students study for exams. AI bots used as psychoeducational tools have a strong potential to enhance the training of future clinicians.

Supervision  

AI-assisted supervisory tools intended to strengthen the student-supervisor relationship, while still nascent, are capable of analyzing session recordings and transcripts, delivering structured feedback to trainees on communication patterns, micro-skills, and competency-linked performance indicators (Louie et al., 2026; Crofford et al., 2026; Cabrera Lozoya et al., 2025). Tools such as Lyssn and mpathic are already in clinical use, trained to identify empathic communication, reflective listening, and open-ended questioning, and can offer trainees alternative phrasing in real time when deviations from a treatment protocol are detected. Platforms such as PsyPilot, designed by psychologists, integrate AI-generated dashboards that track therapeutic alliance, treatment adherence, and symptom trajectories across sessions, providing supervisors with more nuanced, data-informed information (Roca et al., 2026). While Louie et al. (2026) do show a comparable effect with human supervisors, it is important to note the results were compared to another study. Such results likely provide more evidence for the plausibility of such a tool being effective than to direct evidence of its success. In fact, Louie et al. (2026) report that the practice-only group showed worsened empathy than the no-intervention group and an inflated sense of confidence.

The benefits of AI in clinical training extend into the supervisory relationship as well. Clinical supervision currently depends heavily on what trainees remember or choose to report from their sessions. This approach is constrained by individual integrity, memory, selective attention, and the power dynamics that shape the trainee-supervisor relationship (Ciesielski, 2026). AI tools that generate structured transcripts and competency-linked performance ratings from virtual patient interactions and live sessions can transform the supervisor-supervisee dyad. Supervisors can gain access to nuanced behavioral data on trainees’ performance that does not solely rely on secondhand accounts or selective recall. Some benefits include improved accountability, greater specificity in feedback, and a reduced risk that significant skill gaps go unaddressed because they were never surfaced in supervision.

Curriculum Planning and Implementation 

In addition, Gen AI tools such as ChatGPT or Claude show promising results in helping professors, teachers, and trainers develop curricula, lesson plans, quizzes, and exam guides for students (Weightman et al., 2025; Lee et al., 2025; Prégent et al., 2025). Lee et al. (2025) note that LLMs have been used successfully to create Script Concordance Tests (SCTs), which help refine clinical reasoning under conditions of uncertainty. Prégent et al. (2025) identified using AI for educational content, including clinical vignettes, summaries, quizzes, and exam preparation materials, as a common application of AI in psychiatry and psychology training. Similarly, Weightman et al. (2025) reported uses for AI-generated material in group tutorials and, in some cases, for full curriculum design.

Research Support 

 Another area in which AI can be utilized is finding funding sources and grants, assisting with data collection and analysis, and disseminating information (Crofford et al., 2026). In their qualitative study of counselor educators, Crofford et al. (2026) found that participants saw general-purpose AI tools as helpful across the research process itself, pointing to platforms such as Litmaps and Scholarcy for identifying gaps in the literature and even assist in drafting literature reviews. Prégent et al. (2025) also identified the potential for AI to support program and policy development at the institutional level.

Case Records and Documentation  

AI-generated clinical documentation is perhaps the most widely adopted application among practicing professionals, and the proliferation of electronic health record (EHR) systems makes avoiding AI-enhanced documentation nearly impossible for trainees in many settings.

Fang et al. (2025) comment that general-purpose tools such as ChatGPT, when tasked with creating a psychiatric assessment, can produce a seemingly reasonable output. More advanced systems in current use, such as Nuance DAX (now Microsoft Dragon Copilot) , Lyssn, PsyPilot, and SimplePractice, to name a few, use audio/video recording features that automatically transcribe and structure session content into clinical notes with minimal clinician input. For trainees and practicing clinicians, the substantial reduction in administrative burden these tools offer is significant; and if used correctly, they can also identify areas for trainee improvement.

Testing

AI systems are increasingly used to score psychological tests, generate interpretive narratives, and even produce full draft reports in a fraction of the time required for human-generated output. Computerized scoring and computer-generated reports are not new in psychological assessment; however, the sophistication of these tools has increased substantially with the advent of GenAI. Tools currently in use, such as Psynth, Psychreport.ai, and AI Report Writer (PAR), generate structured diagnostic reports from raw data and provide interpretive narratives that clinicians can review and modify before finalization. These tools are positioned as assistive aids that streamline clinical workflow with meaningful efficiency gains.

The data on using AI for psychological testing is also promising, showing accuracy and a significant reduction in time spent on report writing. One study found that AI-generated (ChatGPT-4) psychological reports were completed in approximately 91 seconds, compared to roughly 2.5 hours for human-written reports (Lockwood et al., 2025). Furthermore, the study noted that human raters (licensed psychologists) favored the A. I.’s suggestions in the suggestions section, but preferred the writing style, organization, and overall quality of the human-produced reports.  While these preliminary results show that a tool like this can potentially save hours of work and produce more timely reports for patients, there are considerable weaknesses to address. 

Risks and Concerns

There are important methodological and conceptual concerns regarding the current research on A.I. and clinical training. Kilincel et al. (2026) noted that the studies in their meta-analysis differed considerably in how they were designed, who took part, and what was measured, what kind of AI was being used, and that a large share used small samples, exploratory pilot formats, or before-and-after comparisons that lacked a control group. These limitations constrain causal attribution and make it difficult to conclude efficacy. Most existing research relies on self-report or trainee perceptions collected immediately after using an AI tool, with virtually no evidence on what happens when those trainees subsequently sit with real patients.

Deskilling Cognitive Offloading

Deskilling.  It is worth naming the most consequential overarching risk AI poses to clinical training, one that is not unique to any single tool or application but can operate across all of them. That risk is cognitive offloading and deskilling–the process by which trainees either fail to develop foundational clinical skills in the first place because AI performs those functions for them, or because they gradually lose skills they have already developed through sustained reliance on automated scaffolding. Braverman (1974) first named the concept of deskilling in the context of industrial labor, describing how the concentration of expertise in technological systems systematically eroded workers’ own competencies. The mechanism is directly applicable here. A recent study (Gerlich, 2025) surveyed and interviewed 666 people spanning a wide range of ages and levels of education and found that heavier AI use was associated with weaker critical thinking, an effect explained by cognitive offloading. Furthermore, the study found that younger participants who relied more on AI had the lowest critical thinking scores (Gerlich, 2025). One of AI’s most significant threats to clinical training is the risk of producing a generation of clinicians whose critical thinking and reasoning skills have been unintentionally deprioritized, leading to the advancement of psychologists who lack independent clinical reasoning (van Zyl, 2026).

Cognitive Offloading. The mechanism by which deskilling operates in this context is cognitive offloading, by using external tools to lighten the cognitive demands of a task (Risko & Gilbert, 2016; Gerlich, 2025; van Zyl, 2026). Offloading is not inherently problematic; it is, in fact, often the point of a well-designed tool. However, Grinschgl and colleagues (2021) provide experimental evidence that cognitive offloading impairs memory when learners do not hold an explicit goal to encode what is being offloaded. The training implications are direct: a trainee who uses AI to synthesize and generate a case formulation or a session note, and then reviews it, is engaged in a qualitatively different cognitive task than one who generates the formulation or note independently. Those using tools to offload tasks may perform better in the moment but struggle with memory performance (Grinschgl et al., 2021). However, the same authors note that participants who were aware they would need to recall the offloaded material showed improved memory performance, suggesting ways to mitigate some of the negative effects. It is important to distinguish AI as a tool for performance and efficiency from AI as a tool for deep learning, which should be prioritized in training settings.

Because “doing is learning”, it can be difficult to assess the cognitive-learning impact of tools that make life more efficient for students. Efficiency, which by definition, implies doing less, and saving time is understandably valued by any busy student or professor. However, that efficiency comes at a cost. For example, for the practicing professional, documentation may feel like an administrative task that could benefit from efficiency tools. For trainees, though, the implications are more complicated. Documentation is itself a clinical skill that teaches critical thinking, and trainees who produce notes primarily through AI-assisted drafting may not develop the habit of structured clinical reflection that the documentation process has traditionally scaffolded (Fang et al., 2025). With the proliferation of EHR tools, how programs address these concerns is critical.

Algorithmic Bias

Algorithmic bias is a concern that spans diverse AI applications. The American Psychological Association (2017, 2021a, 2021b) establishes cultural and structural responsiveness as a foundational competency that spans all aspects of a psychologist’s work, yet LLMs risk oversimplifying client identities and structural conditions and, worse, potentially reinforcing stereotyped presentations. AI systems are trained on out-of-date, non-representative data, perpetuating historical inequities (Ali et al., 2025; Erdemir & Sumbas, 2026). Emerging research suggests that AI systems trained predominantly on non-representative datasets, drawn largely from Western, English-speaking, and often white, educated populations, perform significantly less well for individuals from ethnic minorities and non-native speakers (Erdemir & Sumbas, 2026; Roca et al., 2026; Ali et al., 2025). Programs adopting AI tools in training contexts have a professional obligation to evaluate each tool explicitly for cultural representativeness before adoption. Unfortunately, most of these AI tools are essentially black boxes, leaving programs and clinicians with little basis for evaluating the robustness of an AI’s training data.

Therapeutic Alliance

One serious limitation of AI bots as a training modality is their inability to replicate the relational conditions in which therapeutic competence is actually developed. The quality of the therapeutic alliance is among the most researched and empirically supported predictors of clinical outcome (Bordin, 1979; Flückiger et al., 2018; Safran et al., 2011). Although ruptures and repairs are considered an inevitable and important aspect of the therapeutic relationship (Eubanks et al., 2018), the complexity of an alliance, along with ruptures and repairs, is not easily recreated. VPs do not arrive in crisis outside of session hours, deteriorate unexpectedly, lie, withhold, or resist care in the ways that real patients can and do. However, tolerating those experiences and navigating a real rupture are important aspects of clinical competency.

The Role of Emotion

The emotional dimensions of clinical work (and in particular with regard to the therapeutic alliance) are not skills that can be adequately gained in a simulation environment where the trainee knows, at some level, that there are no real consequences. At some point, every trainee finds themselves worried about a patient’s safety with limited resources available to do anything. The capacity to remain regulated and effective in genuinely high-stakes clinical encounters develops only through exposure to those same high-stakes encounters. This is a structural limitation of AI simulation at present, and it marks a boundary on how much of clinical training can be appropriately delegated to virtual environments.

Supervision

A related concern is what an increasing reliance on AI-generated performance data does to the supervisory relationship itself. Clinical supervision serves a broad range of purposes beyond the evaluative. How AI tools may augment that relationship, or gradually displace it, carries significant implications for how trainees are formed as clinicians. AI may also muddle a supervisor’s ability to think freely and openly; once AI begins making suggestions, it may not always be easy for a supervisor to know how to proceed independently. It is not unreasonable to think that, over time, supervisors will come to rely more on AI output than on their own clinical judgment.

Ciesielski (2026) discusses how AI-generated reports may increase student anxiety when used in formal evaluations of their work. The supervisory relationship is not only a mechanism for skill development; it is itself a formative relational experience that models the kind of attentive, accountable human engagement that clinical work also demands. AI tools positioned as enhancements to supervision may strengthen the relationship by providing concrete behavioral data that supervisors can use for feedback. Even so, the line between support and substitution may not be so easily navigated. Please see the section on Supervision for a more detailed analysis.

Academic Integrity

A concern that has received increasing attention in clinical training contexts is academic integrity. The same generative AI capabilities that make these tools potentially powerful for learning also enable them to complete many of the assessments traditionally used to evaluate trainee competency. ChatGPT-4 has been shown to achieve passing grades on German-language multiple-choice exams, scoring above 90% correct, and, with somewhat lower accuracy, achieved passing grades on Mandarin-language assessments as well (Weightman et al., 2025). The capabilities of these programs will only improve, and results like these have serious implications for how competency is assessed in clinical settings. Programs will need to become more creative in how they assess students’ proficiencies.

Socialization and Professional Identity

Finally, training is not only skill acquisition; it is socialization into a professional role, and AI tools optimized for behavioral feedback may inadvertently narrow what trainees come to understand clinical identity and competence to mean. The formation of professional identity in clinical psychology unfolds through relational and reflective experiences over time. It is, in part, through genuine clinical experiences and supervision relationships that model clinical and ethical reasoning (Rønnestad & Skovholt, 2003).

We risk creating a generation of clinicians more attuned to optimization, and to what can be and is being measured, as the only things that matter in the clinical setting. Training programs that rely on AI-generated performance feedback or synthesis must be attentive to the risk that this feedback may gradually narrow a trainee’s identity, purpose, and understanding of the clinical setting. Hanley (2025) remarks that “if therapy becomes overly mechanised, there is a risk of shifting from being with to merely doing to clients.” Such a shift would alter the field’s foundation.

AI, used as a scaffold that builds toward independent reasoning, critical thinking, and deep learning, has genuine potential to strengthen training in ways to which the field has not previously had access. Used poorly, or too early, or without adequate supervision, it risks producing a generation of practitioners too focused on optimization and measurable outcomes, whose clinical reasoning may become dependent on AI outputs, and even worse, on outputs they are not equipped to evaluate critically. Ciesielski (2026) notes that, as trainees come into greater contact with AI, guidance will be a critical component of their training. Trainees are poorly positioned to independently critically evaluate AI output, which risks reinforcing bad habits or disseminating incorrect information.

Recommendations

There is strong consensus among researchers that AI literacy should be integrated into the core curriculum for clinician education and training (Prégent et al., 2025; Weightman et al., 2025; van Zyl, 2026; Crofford et al., 2026; Ali et al., 2025). Training programs will need committees to develop clear guidelines for using AI appropriately.

AI literacy must not be treated as a secondary concern, but as a foundational core competency moving forward.  This starts with ensuring that all staff are trained on the tools they use and teach, as well as the tools students may have access to, such as LLMs like ChatGPT and Claude. Prégent et al. (2025) suggest interdisciplinary collaboration among relevant disciplines to create meaningful committees and task forces capable of addressing the complexity that no single discipline can handle alone. Such committees can address appropriate adoption of technologies, academic integrity, and clear guidelines around expectations and appropriate use. Ciesielski (2026) notes that the APA’s AI Evaluation Checklist is a useful tool for practitioners to consult when adopting an AI tool, and this will be a strong resource for programs as well.

Students should be informed about how AI tools are developed, what risks and benefits they bring, how they impact learning, and what best practices look like. There should be active engagement in teaching students to think critically about the technologies they use so that they can evaluate, assess, and identify errors. This must include understanding, identifying, and working with the biases inherent in AI tools, including socioeconomic and cultural biases, how data are collected and stored, and the privacy issues that arise (Ali et al., 2025). Furthermore, Erdemir and Sumbas (2026) note that it is prudent for programs to address what to do when a trainee disagrees with an AI output or suggestion. For a beginner who may not have sufficient background to assess the issue, what are the appropriate steps to resolve the disagreement?

Given concerns about cognitive offloading and its potential side effects (Gerlich, 2025), it will be prudent for programs to consider introducing different technologies at different stages of students’ or trainees’ training (van Zyl, 2026). This starts with ensuring students have a strong grasp of a task or required competency before engaging with a technology that may reduce their real or perceived need to think critically. It may also be beneficial to build in tech blackout periods during which students practice skills without using AI. As documentation becomes increasingly automated, students should be required to periodically complete a case conceptualization, case formulation, or testing report without the aid of AI tools.

In sum, student proficiency should be evaluated through a critical lens that allows students to truly demonstrate their level of understanding. In this regard oral exams and direct observation are likely to remain effective means of assessment moving forward.

Use of AI Tools

As LLM-based virtual patients (VPs) become more prevalent and capable, it will be key for programs to treat these tools as potentially useful supplements, but not sufficient for trainees’ exposure to a clinical population (Imam Hossain et al., 2024). Access to a wide range of patients will remain a concern, and VPs should not be allowed to justify becoming lax about trainees’ exposure to real-life patients. In time, licensing boards will likely have to wrestle with how (if at all) virtual patient hours should count toward licensure. It will be critical that this number remain limited to prevent programs and training sites from becoming overly reliant on bots.

The impact of AI on supervision is likely to be significant over time. Used as a supportive tool, it may help deepen the work between supervisor and supervisee; but it should never replace it. The interpersonal component of supervision is as critical. Just as VPs are not a sufficient training tools on their own, neither are virtual supervisors. It will be critical that all parties are clear about the role AI will play in how the relationship functions and whether the student will be assessed. Transparency will be essential. There is still a great deal unknown about the influence AI will have on a supervisor’s ability to think critically and independently. One possible way to maintain independence while still collaborating with AI outputs is to hold some supervisory sessions with it and some without, or to introduce AI’s output only at the end of a session. It is hard to imagine that a supervisor who has read an AI output before a session would not be influenced by it in some way.

Conclusion

Clinical programs with the capacity to do so should prioritize AI research in the clinical setting. Large longitudinal studies assessing impacts and outcomes are certainly needed (Erdemir & Sumbas, 2026; Imam Hossain et al., 2024). Ideally, programs might collaborate to build a sufficient sample size to draw meaningful conclusions about AI’s impact on programs–their instructors, supervisors, trainees, and patients. It will be imperative to understand how AI bots used in training truly translate to the real-world clinical setting.

Overall, AI has significant potential to be a powerful, supportive tool that enhances learning and education for both students and faculty when used correctly. It also carries real risks that, if left unchecked, could have a detrimental impact on the field of psychotherapy. Assessing any AI tool in the context of education and training begins with a foundational question: is the tool integrated as a supportive tool that bolsters learning, retention, and professional judgment, or is it simply a replacement for existing (and effective) clinical education and training?

References: Education & Training

Ali, M., Ali, S., Abbas, Q., Abbas, Z., & Lee, S. W. (2025). Artificial intelligence for mental health: A narrative review of applications, challenges, and future directions in digital health. DIGITAL HEALTH, 11, 1-25. https://doi.org/10.1177/20552076251395548

American Psychological Association. (2017). Multicultural guidelines: An ecological approach to context, identity, and intersectionality. https://www.apa.org/about/policy/multicultural-guidelines.pdf

Bordin, E. S. (1979). The generalizability of the psychoanalytic concept of the working alliance. Psychotherapy: Theory, Research & Practice, 16(3), 252-260. https://doi.org/10.1037/h0085885

Braverman, H. (1974). Labor and monopoly capital: The degradation of work in the twentieth century. Monthly Review Press.

Cabrera Lozoya, D., Conway, M., De Duro, E. S., & D’Alfonso, S. (2025). Leveraging large language models for simulated psychotherapy client interactions: Development and usability study of Client101. JMIR Medical Education, 11, Article e68056. https://doi.org/10.2196/68056

Ciesielski, H. (2026). Ethical considerations for AI in psychological assessment practice and training. Assessment. Advance online publication. https://doi.org/10.1177/10731911261453692

Crofford, H., Bor, E., & Kemer, G. (2026). Counseling professionals’ perspectives on AI integration in education and supervision: A concept mapping study. Counselor Education & Supervision, 65(1), 22-32. https://doi.org/10.1002/ceas.70016

Elyoseph, Z., Levi-Belz, Y., Levkovich, I., Haber, Y., Gramaglia, C. M., López Castroman, J., Cecile, H., & Olie, E. (2026). The effectiveness of multilingual AI-based simulator for suicide risk assessment training in improving self-efficacy among young psychiatrists: A pilot study across twenty languages. BMC Psychiatry, 26, Article 98. https://doi.org/10.1186/s12888-025-07737-9

Erdemir, N., & Sumbas, E. (2026). Integrating artificial intelligence into psychological counseling: A narrative review and governance framework. INQUIRY: The Journal of Health Care Organization, Provision, and Financing, 63, 1-16. https://doi.org/10.1177/00469580261438322

Eubanks, C. F., Burckell, L. A., & Goldfried, M. R. (2018). Clinical consensus strategies to repair ruptures in the therapeutic alliance. Journal of Psychotherapy Integration, 28(1), 60–76. https://doi.org/10.1037/int0000097

Fang, A., Kramer, E. N., & Velicu, V. I. (2025, June). ChatGPT and psychiatric documentation: Balancing trainee education and administrative burden. American Journal of Psychiatry Residents’ Journal, 7-8.

Flückiger, C., Del Re, A. C., Wampold, B. E., & Horvath, A. O. (2018). The alliance in adult psychotherapy: A meta-analytic synthesis. Psychotherapy, 55(4), 316-340. https://doi.org/10.1037/pst0000172

García-Torres, D., Fernández, C., Mira, J. J., Morales, A., & Vicente, M. A. (2025). Using AI-based virtual simulated patients for training in psychopathological interviewing: Cross-sectional observational study. JMIR Medical Education, 11, Article e78857. https://doi.org/10.2196/78857

Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), Article 6. https://doi.org/10.3390/soc15010006

Grinschgl, S., Papenmeier, F., & Meyerhoff, H. S. (2021). Consequences of cognitive offloading: Boosting performance but diminishing memory. Quarterly Journal of Experimental Psychology, 74(9), 1477-1496. https://doi.org/10.1177/17470218211008060

Hanley, T. (2025). Artificial intelligence and the changing landscape of therapy. Counselling and Psychotherapy Research, 25, Article e70056. https://doi.org/10.1002/capr.70056

Hsieh, L.-H., Liao, W.-C., & Liu, E.-Y. (2024). Feasibility assessment of using ChatGPT for training case conceptualization skills in psychological counseling. Computers in Human Behavior: Artificial Humans, 2, Article 100083. https://doi.org/10.1016/j.chbah.2024.100083

Imam Hossain, S., Kelson, J., & Morrison, B. (2024). The use of virtual patient simulations in psychology: A scoping review. Australasian Journal of Educational Technology, 40(6), 76-91. https://doi.org/10.14742/ajet.9559

Kılıncel, S., Bulut, F., Goksel, P., Usta, M. B., Mutluer, T., & Kilincel, O. (2026). Effectiveness of AI-enhanced virtual patients for psychiatric interview training in health professions education: A meta-analysis. Frontiers in Medicine, 13, Article 1834636. https://doi.org/10.3389/fmed.2026.1834636

Lee, Q. Y., Chen, M., Ong, C. W., & Ho, C. S. H. (2025). The role of generative artificial intelligence in psychiatric education – a scoping review. BMC Medical Education, 25, Article 438. https://doi.org/10.1186/s12909-025-07026-9

Lockwood, A. B., Farmer, R. L., Shergill, G., Benson, N. F., & Gilbert, K. (2025). Human vs. machine: Comparing AI-generated and human-written psychological reports. Journal of Psychoeducational Assessment, 43(6), 559–573. https://doi.org/10.1177/07342829251346623

Louie, R., Shah, R. S., Orney, I. H., Pacheco, J. P., Brunskill, E., & Yang, D. (2026). Can LLM-simulated practice and feedback upskill human counselors? A randomized study with 90+ novice counselors. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). ACM. https://doi.org/10.1145/3772318.3791821

O’Brien, K. H. M., Fuxman, S., Humm, L., Tirone, N., Pires, W. J., Cole, A., & Goldstein Grumet, J. (2019). Suicide risk assessment training using an online virtual patient simulation. mHealth, 5, Article 31. https://doi.org/10.21037/mhealth.2019.08.03

Prégent, J., Chung, V.-H.-A., El Adib, I., Désilets, M., & Hudon, A. (2025). Applications of artificial intelligence in psychiatry and psychology education: Scoping review. JMIR Medical Education, 11, Article e75238. https://doi.org/10.2196/75238

Richard, O., Jollant, F., Billon, G., Attoe, C., Vodovar, D., & Piot, M.-A. (2023). Simulation training in suicide risk assessment and intervention: A systematic review and meta-analysis. Medical Education Online, 28(1), Article 2199469. https://doi.org/10.1080/10872981.2023.2199469

Risko, E. F., & Gilbert, S. J. (2016). Cognitive offloading. Trends in Cognitive Sciences, 20(9), 676-688. https://doi.org/10.1016/j.tics.2016.07.002

Roca, P., Zangri, R. M., Rodriguez-Fernandez, G., Sanchez-Pedreño, M., & García del Valle, E. P. (2026). Artificial intelligence in the psychologist’s toolkit: Psypilot as a case study. Frontiers in Psychology, 17, Article 1775464. https://doi.org/10.3389/fpsyg.2026.1775464

Rønnestad, M. H., & Skovholt, T. M. (2003). The journey of the counselor and therapist: Research findings and perspectives on professional development. Journal of Career Development, 30(1), 5-44. https://doi.org/10.1023/A:1025173508081

Safran, J. D., Muran, J. C., & Eubanks-Carter, C. (2011). Repairing alliance ruptures. Psychotherapy, 48(1), 80–87. https://doi.org/10.1037/a0022140

Sarli, G., Rogers, M. L., Bloch-Elkouby, S., Lawrence, O. C., Gomes de Siqueira, A., Yao, H., Lok, B., Foster, A., & Galynker, I. (2024). Using virtual patients to assess and improve clinicians’ emotional self-awareness: A randomized controlled study. Academic Psychiatry, 48, 18-28. https://doi.org/10.1007/s40596-023-01909-z

van Zyl, L. E. (2026). The unintended negative consequences of artificial intelligence use for psychologists. Frontiers in Psychology, 17, Article 1729050. https://doi.org/10.3389/fpsyg.2026.1729050

Wang, R., Milani, S., Chiu, J. C., Zhi, J., Eack, S. M., Labrum, T., Murphy, S. M., Jones, N., Hardy, K., Shen, H., Fang, F., & Chen, Z. Z. (2024). PATIENT-Ψ: Using large language models to simulate patients for training mental health professionals (arXiv:2405.19660). arXiv. https://doi.org/10.48550/arXiv.2405.19660

Weightman, M. J., Chur-Hansen, A., & Clark, S. R. (2025). AI in psychiatric education and training from 2016 to 2024: Scoping review of trends. JMIR Medical Education, 11, Article e81517. https://doi.org/10.2196/81517

Weizenbaum, J. (1966). ELIZA—A computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1), 36–45. https://doi.org/10.1145/365153.365168

Xian, X., Chang, A., Xiang, Y.-T., & Liu, M. T. (2024). Debate and dilemmas regarding generative AI in mental health care: Scoping review. Interactive Journal of Medical Research, 13, Article e53672. https://doi.org/10.2196/53672   

Companies/Products – AI and Psychotherapy Education & Training

Litmaps. (n.d.). Litmaps. Retrieved August 16, 2026, from https://www.litmaps.com/

Lyssn.io, Inc. (n.d.). Lyssn. Retrieved August 16, 2026, from https://www.lyssn.io/

Microsoft. (n.d.). Microsoft Dragon Copilot. Retrieved August 16, 2026, from https://www.microsoft.com/en-us/health-solutions/clinical-workflow/dragon-copilot

mpathic. (n.d.). mpathic. Retrieved August 16, 2026, from https://mpathic.ai/

PAR, Inc. (n.d.). AI Report Writer. Retrieved August 16, 2026, from https://www.parinc.com/products/ai-report-writer

PsychReport.ai. (n.d.). PsychReport.ai. Retrieved August 16, 2026, from https://www.psychreport.ai/

Psynth, Inc. (n.d.). Psynth. Retrieved August 16, 2026, from https://www.psynth.ai/

Roca, P., Zangri, R. M., Rodriguez-Fernandez, G., Sanchez-Pedreño, M., & García del Valle, E. P. (2026). Artificial intelligence in the psychologist’s toolkit: Psypilot as a case study. Frontiers in Psychology, 17, 1775464. https://doi.org/10.3389/fpsyg.2026.1775464

Scholarcy. (n.d.). Scholarcy. Retrieved August 16, 2026, from https://www.scholarcy.com/

SimplePractice. (n.d.). AI-powered Note Taker. Retrieved August 16, 2026, from https://www.simplepractice.com/features/ai-therapy-notes-taker/

AI in Psychotherapy Research

Xu Li and Wonjin Sim

Artificial intelligence (AI) may enhance psychotherapy research primarily by strengthening the scientific infrastructure through which psychotherapy knowledge is synthesized, measured, and advanced. Across multiple domains, recent evidence suggests that machine learning (ML), natural language processing (NLP), large language models (LLMs), and multimodal computational systems are increasingly contributing to psychotherapy research by improving the efficiency of knowledge synthesis, expanding precision in process and outcome measurement, and advancing personalized frameworks for understanding “what works, for whom, when” (Aafjes-van Doorn et al., 2021). Importantly, current evidence does not suggest that AI should replace human scientific or clinical judgment. Rather, its primary value appears to lie in augmenting psychotherapy research capacity while remaining embedded within rigorous human oversight.

Opportunities: Expanding the Scope, Speed, and Precision of Psychotherapy Research

One major area of AI application in psychotherapy research is scientific efficiency, particularly through literature reviews and evidence synthesis. Across healthcare, ecology, information systems, and qualitative research, AI is increasingly used to support title/abstract screening, search strategy refinement, study selection, data extraction, and evidence synthesis through tools such as ML classifiers, active learning systems, NLP models, and LLMs (Blaizot et al., 2022; Marshall & Wallace, 2019; Van Dijk et al., 2023; Fabiano et al., 2024; Wagner et al., 2021). AI-assisted screening systems such as ASReview, RobotSearch, and TrialMind have demonstrated substantial reductions in reviewer burden, often reducing manual screening demands by over 50% while maintaining high recall of relevant studies (Abogunrin et al., 2025; Spillias et al., 2024; Bolanos et al., 2024; Wang et al., 2024). AI can also assist with search string refinement, article recommendation, topic modeling, structured data extraction, and synthesis preparation (Ge et al., 2024; Sušnjak et al., 2024). These developments suggest that AI may help psychotherapy researchers manage increasingly unsustainable literature volumes, reduce labor costs, and accelerate hypothesis generation, particularly in rapidly expanding or under-resourced scientific areas (Berger‐Tal et al., 2024; Thomas et al., 2024; Choi, 2025).

A second major domain involves scientific precision through psychotherapy process and outcome measurement. AI is increasingly used to analyze psychotherapy transcripts, large-scale digital treatment logs, and multimodal datasets incorporating text, audio, and video. NLP and ML systems are now commonly applied to classify or predict alliance, therapist adherence, emotional tone, therapist skills, treatment response, and symptom trajectories from psychotherapy sessions (Goldberg et al., 2020; Malgaroli et al., 2023; Laricheva et al., 2024; Beg et al., 2024). For example, Goldberg et al. (2020) demonstrated that linguistic features from 1,235 psychotherapy sessions modestly predicted client-rated therapeutic alliance (ρ ≈ .15), while Eberhardt et al. (2024) found that transformer-based (an advanced NLP architecture) sentiment analyses captured emotional process variables linked to session dynamics and treatment outcomes. Lalk et al. (2025) further demonstrated that fine-tuned LLMs can identify multiple emotional categories across psychotherapy sessions, with aggregated emotional profiles predicting symptom severity and moderately predicting alliance. Beyond transcript analysis, multimodal systems integrating text, audio, and video have shown stronger alliance prediction (Aafjes-van Doorn et al., 2025), while computational methods for facial expression, posture, prosody, and synchrony are increasingly being explored (Hau et al., 2025). AI is also being applied to large online psychotherapy datasets to study dropout, adherence, provider characteristics, and engagement trajectories (Gutiérrez et al., 2024). Collectively, these developments suggest that AI may substantially improve scalable process measurement, reduce reliance on burdensome self-report or human coding systems, and enable more ecologically valid psychotherapy research.

A third major domain involves scientific advancement through personalized psychotherapy frameworks. Psychotherapy research is increasingly shifting away from identifying universally superior treatments and toward precision-oriented questions concerning which interventions work best for which patients under which conditions. Within this framework, AI and ML are being used to predict treatment response versus nonresponse, stratify patient risk, estimate heterogeneous treatment effects, and support dynamic treatment adaptation (Curtiss & DiPietro, 2025; Jankowsky et al., 2023; Chekroud et al., 2021). Across emotional disorders, ML models have shown good discrimination in predicting treatment response (mean accuracy ≈ .76; AUC ≈ .80), suggesting potential utility for more individualized treatment planning (Curtiss & DiPietro, 2025).

Broader psychiatric applications have also demonstrated that ML can help predict whether patients may benefit more from CBT, interpersonal therapy, medication, or varying treatment intensities (Chekroud et al., 2021). Causal machine learning and meta-learning methods further allow psychotherapy researchers to estimate conditional average treatment effects and move beyond average outcomes toward individualized treatment rules (Künzel et al., 2017; Feuerriegel et al., 2024; Salditt et al., 2023). Dynamic AI systems embedded in digital interventions may additionally allow for near real-time personalization by adapting intervention timing, modality, or intensity based on symptom trajectories and contextual factors (Gual-Montolio et al., 2022; Gutiérrez et al., 2024). Together, these approaches may substantially strengthen psychotherapy science’s capacity to study individualized pathways of change.

Concerns: Methodological Constraints, Ethical Concerns, and Epistemic Caution

Despite substantial promise, the literature consistently emphasizes that AI in psychotherapy research remains constrained by important methodological, ethical, and epistemic limitations.

In evidence synthesis, AI tools remain vulnerable to omission errors, hallucinations, weak reproducibility, and variability in usability and transparency (Clark et al., 2025; Ge et al., 2024; Sušnjak et al., 2024). While AI-assisted screening performs relatively well, standalone generative AI search performs poorly; for example, Clark et al. (2025) found that General AI search missed a median of 91% of target studies. In addition, hallucinations may occur when AI systems generate plausible but incorrect information, such as fabricating references, misreporting study findings, or inventing methodological details that are not part of the original studies, potentially introducing false evidence into reviews. Scholars also caution against over-automation, reduced researcher learning, algorithmic filtering, and biased shaping of what knowledge is surfaced (Foley et al., 2025).

In psychotherapy process measurement, most AI studies remain in relatively early proof-of-concept stages (Aafjes-van Doorn et al., 2021). Predictive performance for complex psychotherapy constructs such as alliance and emotion, while often statistically significant, remains modest (Goldberg et al., 2020; Lalk et al., 2025). Reviews also consistently identify limited linguistic and demographic diversity, weak external validation, insufficient transparency, omitted preprocessing details, and inadequate explainability (Laricheva et al., 2024; Tornero-Costa et al., 2023). The limited representation of diverse linguistic, cultural, and demographic groups is particularly concerning because AI models trained predominantly on homogeneous datasets may learn patterns that do not generalize to underrepresented populations. Consequently, model predictions may reflect or amplify existing biases toward underrepresented populations. Such disparities may compromise the validity, fairness, and clinical applicability of AI-based psychotherapy process measures.

In addition, privacy concerns are especially salient given the sensitivity of psychotherapy transcripts, audio, and video data (Richards, 2024; Kister et al., 2023). Research that uses AI often requires data transfer to external servers, cloud-based processing, or third-party Application Programming Interfaces, increasing the risk of unauthorized access, data leakage, or secondary use beyond the original consent framework. Even when data are de-identified, advances in re-identification techniques and the richness of psychotherapy narratives may make complete anonymization difficult, heightening the risk of unintended disclosure of highly sensitive clinical information.

In personalized psychotherapy research, although methodological tools such as causal ML and meta-learners are increasingly sophisticated, many predictive systems have not yet demonstrated reliable improvements in patient-level outcomes across real-world settings (Chekroud et al., 2021; Aafjes-van Doorn et al., 2021). External validation remains limited, and implementation may be constrained by therapist skill sets, clinician trust, patient preferences, privacy concerns, and fears regarding the impact of AI on therapeutic relationships (Stade et al., 2024). Moreover, inferring “what works for whom” from observational data requires strong causal assumptions, careful bias control, and robust data quality (Feuerriegel et al., 2024; Salditt et al., 2023).

Across all domains, a broader epistemic concern remains: AI may substantially expand what psychotherapy researchers can measure and predict, but measurement expansion should not be conflated with comprehensive understanding. Psychotherapy is fundamentally relational, contextual, and interpretive, and scholars increasingly warn against reducing psychotherapy to techno-centric or overly “data-dominated” models that privilege computational metrics over human meaning (Richards, 2024).

Recommendations: Current Evidence for Responsible Integration

Across all three domains, the strongest evidence supports a human-centered, researcher-in-the-loop framework for AI integration. AI appears most valuable when used as an assistive infrastructure rather than an autonomous replacement for scientific reasoning.

First, in literature synthesis, AI should be used to augment screening, search refinement, and synthesis preparation while preserving human verification, transparency, and explicit reporting of AI-assisted methods (Fabiano et al., 2024; Blaizot et al., 2022). Emerging governance initiatives such as Responsible AI in Evidence Synthesis and proposed PRISMA adaptations emphasize transparency, accountability, preplanning, and reflexivity (Choi, 2025).

Second, psychotherapy process research requires stronger methodological rigor, including larger and more diverse datasets, better external validation, standardized preprocessing, bias auditing, transparent reporting, and explainability (Laricheva et al., 2024; Tornero-Costa et al., 2023). AI systems should be evaluated not solely on predictive performance but also on fairness, interpretability, and theoretical coherence. In addition, there is a need to develop clear guidelines for culturally responsive use of AI in psychotherapy research, including examining impact of sociocultural context and variation in psychotherapy process, identifying, mitigating, and reporting biases that disproportionately affect underrepresented groups, and promoting practices that prevent the reinforcement of existing inequities in psychotherapy and research. Furthermore, it is imperative to develop comprehensive guidelines and technical safeguards to protect privacy and prevent the leakage or unauthorized disclosure of highly sensitive clinical data. The American Psychological Association could develop ethical guidelines for the use of AI in the collection and coding of psychotherapy data and collaborate with major AI companies to promote ethically responsible practices for handling such data. 

Third, in personalized psychotherapy science, predictive and causal AI systems should be translated cautiously into practice only after demonstrating clinically meaningful improvements in patient outcomes, engagement, and therapeutic alliance across diverse populations (Chekroud et al., 2021). Future research should prioritize prospective validation, implementation science, and careful integration with clinician judgment.

Finally, across all domains, psychotherapy researchers should maintain reflexivity regarding how AI influences not only efficiency but also theory-building, interpretation, and epistemology (Wagner et al., 2021). Responsible integration therefore requires combining technological innovation with psychotherapy’s broader commitments to human complexity, contextual understanding, and ethical care.

Summary

Overall, AI offers substantial promise for psychotherapy research by accelerating literature synthesis, expanding process measurement precision, and advancing personalized psychotherapy frameworks. Current evidence suggests that AI may significantly strengthen psychotherapy research’s scientific infrastructure by improving efficiency, scalability, and methodological sophistication. However, important limitations remain substantial, including modest predictive performance, validation weaknesses, bias, privacy concerns, and epistemic risks. The clearest consensus is therefore that AI should be conceptualized not as a replacement for human scholarship or clinical reasoning, but as a powerful yet fallible collaborator. Used transparently, critically, and within robust human-governed frameworks, AI may substantially expand psychotherapy science while preserving the field’s fundamentally relational, contextual, and human-centered foundations.

Disclosure

Preliminary literature identification and synthesis were assisted by Consensus.ai. Language polishing and editorial support were assisted by ChatGPT (OpenAI). All scholarly contents were independently reviewed and approved by the human authors, who retain full responsibility for their accuracy.

References: Research

Abogunrin, S., Muir, J., Zerbini, C., & Sarri, G. (2025). How much can we save by applying artificial intelligence in evidence synthesis? Results from a pragmatic review to quantify workload efficiencies and cost savings. Frontiers in Pharmacology, 16. https://doi.org/10.3389/fphar.2025.1454245

Aafjes-van Doorn, K., Kamsteeg, C., Bate, J., & Aafjes, M. (2021). A scoping review of machine learning in psychotherapy research. Psychotherapy Research, 31, 92–116. https://doi.org/10.1080/10503307.2020.1808729

Aafjes-van Doorn, K., Cicconet, M., Cohn, J., & Aafjes, M. (2025). Predicting working alliance in psychotherapy: A multi-modal machine learning approach. Psychotherapy Research, 35, 256–270. https://doi.org/10.1080/10503307.2024.2428702

Atkinson, C. (2023). Cheap, quick, and rigorous: Artificial intelligence and the systematic literature review. Social Science Computer Review, 42, 376–393. https://doi.org/10.1177/08944393231196281

Beg, M., Verma, M., V., M., & Verma, M. (2024). Artificial intelligence for psychotherapy: A review of the current state and future directions. Indian Journal of Psychological Medicine, 47, 314–325. https://doi.org/10.1177/02537176241260819

Berger‐Tal, O., Wong, B., Adams, C., Blumstein, D., Candolin, U., Gibson, M., Greggor, A., Lagisz, M., Macura, B., Price, C., Putman, B., Snijders, L., & Nakagawa, S. (2024). Leveraging AI to improve evidence synthesis in conservation. Trends in Ecology & Evolution. https://doi.org/10.1016/j.tree.2024.04.007

Blaizot, A., Veettil, S., Saidoung, P., Moreno-García, C., Wiratunga, N., Aceves-Martins, M., Lai, N., & Chaiyakunapruk, N. (2022). Using artificial intelligence methods for systematic review in health sciences: A systematic review. Research Synthesis Methods, 13, 353–362. https://doi.org/10.1002/jrsm.1553

Bolanos, F., Salatino, A., Osborne, F., & Motta, E. (2024). Artificial intelligence for literature reviews: Opportunities and challenges. Artificial Intelligence Review, 57. https://doi.org/10.1007/s10462-024-10902-3

Chekroud, A., Bondar, J., Delgadillo, J., Doherty, G., Wasil, A., Fokkema, M., Cohen, Z., Belgrave, D., DeRubeis, R., Iniesta, R., Dwyer, D., & Choi, K. (2021). The promise of machine learning in predicting treatment outcomes in psychiatry. World Psychiatry, 20. https://doi.org/10.1002/wps.20882

Choi, M. (2025). Artificial intelligence assisted semi-automation tools using for systematic reviews and guideline development. Journal of Evidence-Based Practice. https://doi.org/10.63528/jebp.2025.00008

Clark, J., Barton, B., Albarqouni, L., Byambasuren, O., Jowsey, T., Keogh, J., Liang, T., Moro, C., O’Neill, H., & Jones, M. (2025). Generative artificial intelligence use in evidence synthesis: A systematic review. Research Synthesis Methods, 16, 601–619. https://doi.org/10.1017/rsm.2025.16

Constantino, M. (2024). Measurement-based matching of patients to psychotherapists’ strengths. Journal of Consulting and Clinical Psychology, 92(6), 327–329. https://doi.org/10.1037/ccp0000897

Curtiss, J., & DiPietro, C. (2025). Machine learning in the prediction of treatment response for emotional disorders: A systematic review and meta-analysis. Clinical Psychology Review, 120, 102593. https://doi.org/10.1016/j.cpr.2025.102593

De La Torre-López, J., Ramírez, A., & Romero, J. (2023). Artificial intelligence to automate the systematic review of scientific literature. Computing, 105, 2171–2194. https://doi.org/10.1007/s00607-023-01181-x

Eberhardt, S., Schaffrath, J., Moggia, D., Schwartz, B., Jaehde, M., Rubel, J., Baur, T., André, E., & Lutz, W. (2024). Decoding emotions: Exploring the validity of sentiment analysis in psychotherapy. Psychotherapy Research, 35, 174–189. https://doi.org/10.1080/10503307.2024.2322522

Fabiano, N., Gupta, A., Bhambra, N., Luu, B., Wong, S., Maaz, M., Fiedorowicz, J., Smith, A., & Solmi, M. (2024). How to optimize the systematic review process using AI tools. JCPP Advances, 4. https://doi.org/10.1002/jcv2.12234

Feuerriegel, S., Frauen, D., Melnychuk, V., Schweisthal, J., Hess, K., Curth, A., Bauer, S., Kilbertus, N., Kohane, I., & Van Der Schaar, M. (2024). Causal machine learning for predicting treatment outcomes. Nature Medicine, 30, 958–968. https://doi.org/10.1038/s41591-024-02902-1

Foley, K., McLean, C., De Zylva, R., Asa, G., Maio, J., Batchelor, S., Dzando, G., & Dimassi, A. (2025). Developing a critical imagination for how researchers can use artificially intelligent tools reflexively and responsibly during qualitative literature reviews. International Journal of Qualitative Methods, 24. https://doi.org/10.1177/16094069251316249

Ge, L., Agrawal, R., Singer, M., Kannapiran, P., De Castro Molina, J., Teow, K., Yap, C., & Abisheganaden, J. (2024). Leveraging artificial intelligence to enhance systematic reviews in health research: Advanced tools and challenges. Systematic Reviews, 13. https://doi.org/10.1186/s13643-024-02682-2

Goldberg, S., Flemotomos, N., Martinez, V., Tanana, M., Kuo, P., Pace, B., Villatte, J., Georgiou, P., Van Epps, J., Imel, Z., Narayanan, S., & Atkins, D. (2020). Machine learning and natural language processing in psychotherapy research: Alliance as example use case. Journal of Counseling Psychology, 67(4), 438–448. https://doi.org/10.1037/cou0000382

Gual-Montolio, P., Jaén, I., Martínez-Borba, V., Castilla, D., & Suso‐Ribera, C. (2022). Using artificial intelligence to enhance ongoing psychological interventions for emotional problems in real- or close to real-time: A systematic review. International Journal of Environmental Research and Public Health, 19. https://doi.org/10.3390/ijerph19137737

Gutiérrez, G., Stephenson, C., Eadie, J., Asadpour, K., & Alavi, N. (2024). Examining the role of AI technology in online mental healthcare: Opportunities, challenges, and implications, a mixed-methods review. Frontiers in Psychiatry, 15. https://doi.org/10.3389/fpsyt.2024.1356773

Hau, S., Rugolon, F., Samuels, T., & Högman, L. (2025). Let’s talk about non-verbal communication: Using AI and machine learning for the investigation of interpersonal psychotherapeutic interactions. The Scandinavian Psychoanalytic Review, 48, 51–62. https://doi.org/10.1080/01062301.2025.2539549

Jankowsky, K., Krakau, L., Schroeders, U., Zwerenz, R., & Beutel, M. (2023). Predicting treatment response using machine learning: A registered report. British Journal of Clinical Psychology. https://doi.org/10.1111/bjc.12452

Kister, K., Laskowski, J., Makarewicz, A., & Tarkowski, J. (2023). Application of artificial intelligence tools in diagnosis and treatment of mental disorders. Current Problems of Psychiatry. https://doi.org/10.12923/2353-8627/2023-0001

Künzel, S., Sekhon, J., Bickel, P., & Yu, B. (2017). Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the National Academy of Sciences of the United States of America, 116, 4156–4165. https://doi.org/10.1073/pnas.1804597116

Lalk, C., Targan, K., Steinbrenner, T., Schaffrath, J., Eberhardt, S., Schwartz, B., Vehlen, A., Lutz, W., & Rubel, J. (2025). Employing large language models for emotion detection in psychotherapy transcripts. Frontiers in Psychiatry, 16. https://doi.org/10.3389/fpsyt.2025.1504306

Laricheva, M., Liu, Y., Shi, E., & Wu, A. (2024). Scoping review on natural language processing applications in counselling and psychotherapy. British Journal of Psychology. https://doi.org/10.1111/bjop.12721

Malgaroli, M., Hull, T., Zech, J., & Althoff, T. (2023). Natural language processing for mental health interventions: A systematic review and research framework. Translational Psychiatry, 13. https://doi.org/10.1038/s41398-023-02592-2

Marshall, I., & Wallace, B. (2019). Toward systematic review automation: A practical guide to using machine learning tools in research synthesis. Systematic Reviews, 8. https://doi.org/10.1186/s13643-019-1074-9

Nguyen-Trung, K., Saeri, A., & Kaufman, S. (2024). Applying ChatGPT and AI-powered tools to accelerate evidence reviews. Human Behavior and Emerging Technologies. https://doi.org/10.1155/2024/8815424

Nie, J., Shao, H., Fan, Y., Shao, Q., You, H., Preindl, M., & Jiang, X. (2024). LLM-based conversational AI therapist for daily functioning screening and psychotherapeutic intervention via everyday smart devices. ACM Transactions on Computing for Healthcare. https://doi.org/10.48550/arxiv.2403.10779

Richards, D. (2024). Artificial intelligence and psychotherapy: A counterpoint. Counselling and Psychotherapy Research. https://doi.org/10.1002/capr.12758

Ruan, M., Fan, J., Liu, M., Meng, Z., Zhang, X., & Zhang, C. (2025). Artificial intelligence for the science of evidence synthesis: How good are AI-powered tools for automatic literature screening? BMC Medical Research Methodology, 25. https://doi.org/10.1186/s12874-025-02644-9

Salditt, M., Eckes, T., & Nestler, S. (2023). A tutorial introduction to heterogeneous treatment effect estimation with meta-learners. Administration and Policy in Mental Health, 51, 650–673. https://doi.org/10.1007/s10488-023-01303-9

Spillias, S., Tuohy, P., Andreotta, M., Annand-Jones, R., Boschetti, F., Cvitanovic, C., Duggan, J., Fulton, E., Karcher, D., Paris, C., Shellock, R., & Trebilco, R. (2024). Human-AI collaboration to identify literature for evidence synthesis. Cell Reports Sustainability. https://doi.org/10.1016/j.crsus.2024.100132

Stade, E., Stirman, S., Ungar, L., Boland, C., Schwartz, H., Yaden, D., Sedoc, J., DeRubeis, R., Willer, R., & Eichstaedt, J. (2024). Large language models could change the future of behavioral healthcare: A proposal for responsible development and evaluation. NPJ Mental Health Research, 3. https://doi.org/10.1038/s44184-024-00056-z

Sušnjak, T., Hwang, P., Reyes, N., Barczak, A., Mcintosh, T., & Ranathunga, S. (2024). Automating research synthesis with domain-specific large language model fine-tuning. ACM Transactions on Knowledge Discovery from Data, 19, 1–39. https://doi.org/10.1145/3715964

Thomas, I., Roche, P., & Grêt-Regamey, A. (2024). Harnessing artificial intelligence for efficient systematic reviews: A case study in ecosystem condition indicators. Ecological Informatics, 83, 102819. https://doi.org/10.1016/j.ecoinf.2024.102819

Tornero-Costa, R., Martínez-Millana, A., Azzopardi-Muscat, N., Lazeri, L., Traver, V., & Novillo-Ortiz, D. (2023). Methodological and quality flaws in the use of artificial intelligence in mental health research: Systematic review. JMIR Mental Health, 10. https://doi.org/10.2196/42045

Van Dijk, S., Brusse-Keizer, M., Bucsán, C., Van Der Palen, J., Doggen, C., & Lenferink, A. (2023). Artificial intelligence in systematic reviews: Promising when appropriately used. BMJ Open, 13. https://doi.org/10.1136/bmjopen-2023-072254

Wagner, G., Lukyanenko, R., & Paré, G. (2021). Artificial intelligence and the conduct of literature reviews. Journal of Information Technology, 37, 209–226. https://doi.org/10.1177/02683962211048201

Wang, Z., Cao, L., Danek, B., Zhang, Y., Jin, Q., Lu, Z., & Sun, J. (2024). Accelerating clinical evidence synthesis with large language models. NPJ Digital Medicine, 8. https://doi.org/10.1038/s41746-025-01840-7

Reflections and Recommendations

Stewart Cooper and John Gavazzi

Four sets of authors examined four domains of professional psychotherapy. They worked independently, drew on largely separate literatures, and reached conclusions that are recognizably compatible. That convergence is itself a result worth stating. What follows draws those conclusions together in four areas: the main concepts that recur across domains, their implications for clinical work, guidelines for responsible use, and the philosophical and ethical foundations underlying both.     

Main Concepts

  1. Human Dignity. Human dignity is fundamental and not reducible to data or algorithms. Psychologists, not the tools, carry the obligation to keep that value at the center of the work. That obligation is realized in ordinary decisions rather than announced in policy statements. A patient who is reduced to a classification becomes diminished regardless of how accurate the classification is.
  2. Criteria for Evaluation. Three criteria govern the evaluation of any AI application: its effect on the therapeutic alliance, its preservation of clinician accountability, and its respect for the patient’s narrative integrity. These criteria apply across domains, ensuring that the same three questions posed to a clinical decision support tool are equally relevant to a supervisory feedback dashboard, an ambient AI scribe, or an automated screening system in a research protocol.
  3. Appraisal of Output. AI tools can support documentation, training, and supervision, but every output requires appraisal for accuracy, embedded cultural assumptions, and bias. Appraisal is bounded by what the reviewer already knows, since one cannot catch what one lacks in the framework to name as missing. This is the rationale for independent, expert command of the material rather than a reader merely checking output for plausibility.
  4. Principal Risks. The principal risks include automation bias, deskilling, cultural bias, and privacy threats; these require cautious, informed application. Improved technology does not resolve any of them. Each risk also accumulates quietly rather than announcing itself, as deskilling arrives one outsourced formulation or note at a time, and a confidentiality exposure is ordinarily discovered after the fact.
  5. Pattern Recognition and Clinical Understanding. AI’s pattern recognition differs from genuine clinical understanding, which involves contextual, cultural, and relational cues interpreted within a particular relationship. No statistical model performs that work. What a model produces is a plausible account of a patient in general, while what the psychologist produces is an account of this patient, at this moment, in this relationship.

Clinical Implications

  1. Reason First. Reason independently before consulting a model. Output that arrives first becomes an anchor, and the psychologist’s own analysis is reduced to a revision of it. Reasoning first also produces something to compare the output against, which converts the model’s contribution from a conclusion into a second opinion.
  2. Verification. Treat AI-generated material as an object of critical analysis rather than as a draft awaiting light editing. Verify any claim about statute, regulation, or code section against a primary source. The same standard applies to the literature, since a plausible reference list is not evidence that the references exist or that they say what the output attributes to them.
  3. Confidentiality. Protect confidentiality through rigorous de-identification, verified vendor agreements for any tool that handles protected health information, and the recognition that some cases cannot be described in enough detail to be useful without becoming identifiable. Re-identification risk rises with the richness of the material, so transcripts, audio, and video warrant considerably more caution than a written summary.
  4. Documentation Review. Review AI-generated documentation specifically for treatment strategy and intervention rationale. Cultural context and embeddedness of oppression-based assumptions have to be considered before the note is generated rather than corrected afterward. Clinicians and trainees who read back their own AI-drafted notes absorb whatever account those notes give, so a symptom-heavy record gradually shapes how the case itself is understood.
  5. Patient Use of AI. Ask patients directly about their own use of general-purpose AI tools for mental health purposes, much as one would inquire about other self-help practices, and use the opening for education about the limitations. The question is ordinarily not whether a patient uses these tools but how, which is why the inquiry belongs in routine practice rather than in the aftermath of a problem.

Guidelines for Responsible Use

  1. Apply the Three Criteria Before Adoption. Evaluate whether a given application enhances or erodes the therapeutic or supervisory relationship, whether it preserves rather than diffuses psychologist accountability, and whether it respects the patient’s narrative integrity and cultural context. The criteria are most useful before adoption, rather than after it has been built into a workflow.
  2. Adjunctive Use. Use AI as an adjunct, with transparency, independent validation, and critical engagement. Adjunctive status is a matter of structure rather than intention. A psychologist who forms her own risk formulation and then reads the model’s output is using an adjunct, while a psychologist who queries the model before reaching a formulation has delegated the judgment, whatever the note says afterward.
  3. Visible Modeling. Model responsible use in supervision and training in order to promote a professional culture of ethical excellence. Modeling requires showing the revisions and the rejected outputs, not just the polished result.

Philosophical and Ethical Foundations

The Work Group recognizes that its members and the Society’s readership hold diverse philosophical and religious commitments. We draw on Pope Leo XIV’s recent encyclical not as a religious authority but as a serious and timely philosophical treatment of human dignity in the age of artificial intelligence. Its central claim is that persons must not be reduced to data, and that technologies must serve rather than supplant genuine human encounters. We welcome engagement with any religious or spiritual tradition that offers thoughtful support for these arguments.

  1. Resist Reductionism. Pope Leo XIV’s recent encyclical on artificial intelligence and the empirical psychotherapy literature both resist reductionism, that is, the flattening of complex and irreducible persons into manageable categories. The convergence carries weight because it arrives from two independent directions, one philosophical and one empirical, and neither depends on the other for its force.
  2. Cultural Shaping. Authenticity and presence are culturally shaped; algorithms cannot replicate nuanced human attunement. A system trained toward the statistical center of its data will reproduce a majority-culture default and present it as a neutral standard.
  3. Case-by-case Judgment. Future AI integration depends on case-by-case clinical judgment, cultural humility, and ongoing professional oversight. No general rule will settle whether a particular tool belongs in a particular case, which places the burden on judgment exercised close to the work rather than on policy written at a distance from it.

A Closing Argument

The report closes with the argument advanced by contributing author John Gavazzi in his online essay, Magnifica Humanitas: Human Dignity, Artificial Intelligence, and the Essence of Psychological Practice. Drawing on Pope Leo XIV’s recent encyclical on AI (2026), Gavazzi (2026) argued that authentic clinical encounters rest on vulnerability, presence, and mutual recognition, none of which a statistical model can replicate. He further argued that what counts as authenticity and presence is defined by the patient rather than supplied by any generic account of therapeutic presence, and that an algorithm cannot read the culturally embedded signals of trust and safety that allow a person to show up as herself in treatment. As he put it, “digital empathy does not equate to human bonding.”

His second claim concerns where these questions get settled. Gavazzi held that the future of ethical AI in mental health will not be determined by engineers or theologians. Instead, it will be worked out incrementally through individual cases, in clinical offices and supervision sessions, by practitioners willing to engage AI critically, to advocate for transparent and validated tools, and to refuse to let efficiency displace the therapeutic relationship. Using AI well develops the way any other clinical competency develops: through an iterative process of doing the work, reflecting on what the work reveals, and adjusting accordingly. Psychologists who supervise, consult, or train carry particular responsibility, since trainees who see only polished output never witness the revision and rejection that produced it. Making that reasoning visible is what turns responsible AI use into a professional expectation rather than an exception. That includes showing where the model’s suggestions fit the client’s context, where they needed reshaping to be useful, and where culturally informed understanding had to supply what the model could not.

Final Thoughts

This report will age. The tools named in it will be superseded, the studies cited will be replaced by better ones, and some of the cautions offered here will look either overdrawn or insufficient within a few years. What will not date is the reason the caution was warranted.

We opened this report with a claim that bears repeating, now not as a principle but as a conclusion. A model can be given every word a patient has spoken and still have nothing at stake in what happens next. The psychologist does, and so does the patient. Every recommendation in this report follows from that disparity, and it is what made this work worth doing before any of these tools existed. And, it is what will make it worth continuing long after the current generation of tools has been replaced.

Our profession does not need to resolve every question about AI technologies in order to act responsibly. We need to hold the line on what cannot be delegated, to evaluate each tool against that fundamental imbalance, and to model to the next generation of psychologists what critical, ethical engagement looks like in practice. That is not a technological challenge. It is a human one.  

References: Reflections

Gavazzi, J. (2026, July 17). Magnifica Humanitas: Human dignity, artificial intelligence, and the essence of psychological practice. Ethics and Psychology. https://www.ethicalpsychology.com/2026/07/magnifica-humanitas-human-dignity.html

Leo XIV. (2026). Magnifica humanitas: On safeguarding the human person in the time of artificial intelligence [Encyclical letter]. Libreria Editrice Vaticana. https://www.vatican.va/content/leo-xiv/en/encyclicals/documents/20260515-magnifica-humanitas.html

Tags

Citation

Society for the Advancement of Psychotherapy. (2026). Artificial Intelligence and Psychotherapy: Opportunities, Challenges, and Recommendations. https://societyforpsychotherapy.org/artificial-intelligence-and-psychotherapy-opportunities-challenges-and-recommendations/

References

No references.

Comments

Be the first to share your thoughts.

Leave a comment

Your email address will not be published. Comments are reviewed before they appear.