728x90
I made a editting tool for speakder diarization database. I modified the origin gryannote version and added some more functions.
The following figure represent the results of Chinese shorts duration of 5.0secs.

The dark gray of wave form represnet background noise( music).
The colored segments represent the exracted speaker utterences.
It is accurrate in segmenting where are a human voices. [ Segmentation (pyannote/segmentation-3.0) ]
If you gather the same colored segments, it has a different voices. So it has a low performance in speaker diarization caused by confusiong voice embedding vectors which was fixed by background noise. [ Embedding (pyannote/wespeaker-voxceleb-resnet34-LM) ]
728x90
반응형
