Initial versions of the algorithm were tested on numerous PSGs, recorded in different laboratories and with different recording equipment. Epochs with differences between the auto-stage and the manual stage were identified in each development PSG. The raw data of said epochs were reviewed by highly qualified, experienced RPSGTs. A decision was made as to whether the differences in stage resulted from manual scoring error, auto-scoring error or that the epoch was equivocal and can legitimately be scored either way (see reference 1 for details of this process). Epochs deemed to reflect auto-scoring errors were collected.
Periodically, these epochs were reviewed by MY and the errors are classified into different categories. Adjustments to the relevant algorithms were made to correct the errors with due consideration to the possible impact of the change on other scoring processes. The PSGs with erroneous epochs were re-scored with the modified algorithms.
When the errors were corrected in these files, a set of 70 reference PSGs scored by 3 senior RPSGTs was re-scored by MSS and the % agreement in each PSG and in the total set was determined. A satisfactory algorithm correction was accepted when the change in the relevant algorithms did not result in deterioration of % agreement of the reference PSGs in any variable. Any such deterioration led to further revision of the relevant algorithms, and the process was repeated until the original errors were corrected without degrading the scoring of the reference PSGs.
As more external PSGs were tested and new errors were found, the same process was repeated. This process continued for several years until we were satisfied that no further improvements could be made.
Relevant References: