Abhijith Sunil

AI Systems Architect · Kerala, India

I build and audit the measurement layer of mental-health AI.

The instruments, pipelines and uncertainty accounting that determine whether a clinical system's outputs mean anything.

0.0 0.2 0.4 0.6 0.8 1.0 CEILING 0.86 NOT REACHABLE r_max = sqrt(0.74) = 0.86 correlation with the construct
Published two-week test–retest reliability for the PHQ-9 sits around 0.74. Spearman's 1904 correction for attenuation bounds the correlation between an observed measure and the construct it stands for at the square root of that reliability — about 0.86. The hatched region is unavailable to any model trained against this instrument, at any scale of data or architecture. The mathematics is a century old. How often it gets reported in clinical machine learning is the question I am currently coding a sample of published papers to answer.

Mathematics taught me what a proof is. Statistics taught me what evidence is, and how much less of it we usually have than we think. Data science taught me how quickly a number becomes a decision. Then I started building systems in mental healthcare and found the uncomfortable version of all three: we were making consequential decisions from instruments whose error nobody had ever written down.

There is no blood test for depression. Every model in this field is trained against a questionnaire, a clinician rating, or a structured interview, and those instruments are noisy in ways that are measurable and mostly unmeasured. My work is one question: how much does this system actually know, and does anyone downstream of it know that?

Work

Three things, and where each of them actually is. Nothing here links to something that does not yet exist.

Measurement debt

Definition in draft

The unpaid validation work a clinical AI system assumes has already been done. A working definition, the six layers between a patient and a prediction, and the four lines of the ledger.

Reliability ceilings in mental-health prediction

In progress

A coded sample of published papers that predict a mental-health outcome, asking a narrow question: is the reliability of the outcome instrument reported, and does the reported performance sit inside the bound that reliability implies? Protocol first, then coding. Built entirely from published literature.

mdledger

Planned

An open-source Python package for the arithmetic this argument keeps needing: attenuation-implied ceilings, standard error of measurement, minimal detectable change, reliable change. One line each, so the check stops being optional.

I am not a clinician. I do not diagnose, treat, or advise on individual care. I make no efficacy claims about any product. Views are my own, and I do not write about client systems, clinical data, or individual cases.

Background

I hold an MTech in Data Science and Analytics from Cochin University of Science and Technology, an MSc in Statistics, and a BSc in Mathematics. I work on clinical assessment platforms, AI-assisted intake, longitudinal outcome measurement, qEEG and LORETA data pipelines, and decision-support infrastructure for mental healthcare. I teach statistics and data science as a guest lecturer and workshop facilitator.

I am also a theatre actor, director and writer, with university-level recognition for acting. I don't write about it here. It is where I learned that a true thing badly delivered does not get heard.

Contact

Worth writing about: research collaboration, a technical disagreement, speaking or teaching, or public-safe measurement and architecture work.

abhijithsunilkpza@gmail.com