Volume 2,Issue 7
基于随机森林与模糊森林的精神障碍高维表型特征筛选与分类研究
精神障碍具有较强的临床异质性,不同疾病类别之间常存在症状重叠和表型交叉。如何从高维表型数据中筛选具有分类意义的关键变量,是精神障碍数据驱动研究中的重要问题。本文以Consortium for Neuropsychiatric Phenomics表型数据集为研究对象,围绕精神分裂症、双相情感障碍和注意缺陷多动障碍三类精神障碍,构建高维表型特征筛选与分类分析框架。针对表型变量数量多、变量间相关性较强等问题,本文分别采用随机森林和模糊森林方法进行特征重要性排序与分类建模,并设置三类疾病分类和单病种识别两类实验场景。结果显示,bprs_positive、adult_totalseverity、ns2、etotal 和mpq_score 等变量在不同分类任务中具有较高重要性。随机森林和模糊森林在重要特征识别方面具有一定一致性,但二者在不同疾病识别场景中的分类表现存在差异。研究结果表明,基于随机森林和模糊森林的特征筛选方法能够为精神障碍高维表型数据分析提供有效工具,也为后续开展精神障碍表型特征识别与分类研究提供了参考。
[1] American Psychiatric Association. Diagnostic criteria from DSM-IV-TR[M]. American Psychiatric Pub, 2000.
[2] Frances A J, Widiger T. Psychiatric diagnosis: Lessons from the DSM-IV past and cautions for the DSM-5 future[J]. Annual Review of Clinical Psychology, 2012, 8: 109–130.
[3] Van Dam N T, et al. Data-driven phenotypic categorization for neurobiological analyses: Beyond DSM-5 labels[J]. Biological Psychiatry, 2017, 81(6): 484–494.
[4] Poldrack R A, et al. A phenome-wide examination of neural and cognitive function[J]. Scientific Data, 2016, 3(1): 1–12.
[5] Breiman L. Random forests[J]. Machine Learning, 2001, 45(1): 5–32.
[6] Qi Y. Random forest for bioinformatics[M]//Ensemble Machine Learning. Springer, 2012: 307–323.
[7] Liaw A, Wiener M. Classification and regression by randomForest[J]. R News, 2002, 2(3): 18–22.
[8] Conn D, et al. Fuzzy forests: Extending random forests for correlated, high-dimensional data[J]. 2015.
[9] Karalunas S L, et al. Subtyping attention-deficit/hyperactivity disorder using temperament dimensions: Toward biologically based nosologic criteria[J]. JAMA Psychiatry,
2014, 71(9): 1015–1024.
[10] Patton J H, Stanford M S, Barratt E S. Factor structure of the Barratt Impulsiveness Scale[J]. Journal of Clinical Psychology, 1995, 51(6): 768–774.
[11] Krueger R F, et al. Externalizing psychopathology in adulthood: A dimensional-spectrum conceptualization and its implications for DSM-V[J]. Journal of Abnormal Psychology, 2005, 114(4): 537.