Yesterday, Friday July 5th I signed off on the stimuli.
----------
Verifying the Data
----------
I've handchecked both the duration and pitch resynthesis output. The pitch resynthesis is working perfectly. The first and last iterations correspond to the input and output and the intermediate interpolations are correctly being generated.
For the duration, the timing appears to be off by a small amount (~.001 second) for the whole file, observing the last iteration. I tried to fix this or find the source of this difference but was unable to. The difference is small enough that I am willing to overlook it for now. I do not believe that it will cause any major problem in our experiment.
----------
Data Quality
----------
Both the duration- and pitch- resynthesized data did not sound very good. The default pitch settings were 75 to 600. I felt that this must be suitable because that is a pretty generous range. For a few of the output files it was not bad but for others, the results were unusable.
For my voice, I tweaked these to 50 to 300. The drop from 75 to 50 was significant. I tried many values around 300 but did not find a significant difference so I left it at 300. This greatly improved the data for the duration resynthesis. I looked at the output for when the resynthesis was word-level and again at phone-level. I expected the phone-level resynthesis would provide better results but the results were surprisingly indistinguishable in many situations although even more surprisingly it was even worse in some. (I hand checked the duration tiers to double check that I wasn't just running the same code in both cases.)
So, finally, I am going with word-level resynthesis. Even still, some of the audio files did not turn out well, so I may need to gather new recordings of similar sentences and then p
As for the pitch resynthesis, the output sounds better but is still tinny. Paul Boersma commented that this can happen if one tries to flatten the pitch but shouldn't happen if changing the kind of pitch contour. I played around with the window size but nothing sounded better than the default. There are no other parameters to the pitch resynthesis, so I don't think there is anything more we can do. The artefacts are noticeable but I think they're not too distracting.
Paul Boersma's discussion of this issue with pitch resynthesis:
http://uk.dir.groups.yahoo.com/group/praat-users/message/5745
----------
Up Next
----------
I need to figure out exactly how we are going to use these stimuli. The idea is to do something like AB or AXB analysis. Subjects try to determine which is a better, closer, or more correct usage of one item or which is more appropriate to use in a given context. Should be able to report on this soon.
----------
Verifying the Data
----------
I've handchecked both the duration and pitch resynthesis output. The pitch resynthesis is working perfectly. The first and last iterations correspond to the input and output and the intermediate interpolations are correctly being generated.
For the duration, the timing appears to be off by a small amount (~.001 second) for the whole file, observing the last iteration. I tried to fix this or find the source of this difference but was unable to. The difference is small enough that I am willing to overlook it for now. I do not believe that it will cause any major problem in our experiment.
----------
Data Quality
----------
Both the duration- and pitch- resynthesized data did not sound very good. The default pitch settings were 75 to 600. I felt that this must be suitable because that is a pretty generous range. For a few of the output files it was not bad but for others, the results were unusable.
For my voice, I tweaked these to 50 to 300. The drop from 75 to 50 was significant. I tried many values around 300 but did not find a significant difference so I left it at 300. This greatly improved the data for the duration resynthesis. I looked at the output for when the resynthesis was word-level and again at phone-level. I expected the phone-level resynthesis would provide better results but the results were surprisingly indistinguishable in many situations although even more surprisingly it was even worse in some. (I hand checked the duration tiers to double check that I wasn't just running the same code in both cases.)
So, finally, I am going with word-level resynthesis. Even still, some of the audio files did not turn out well, so I may need to gather new recordings of similar sentences and then p
As for the pitch resynthesis, the output sounds better but is still tinny. Paul Boersma commented that this can happen if one tries to flatten the pitch but shouldn't happen if changing the kind of pitch contour. I played around with the window size but nothing sounded better than the default. There are no other parameters to the pitch resynthesis, so I don't think there is anything more we can do. The artefacts are noticeable but I think they're not too distracting.
Paul Boersma's discussion of this issue with pitch resynthesis:
http://uk.dir.groups.yahoo.com/group/praat-users/message/5745
----------
Up Next
----------
I need to figure out exactly how we are going to use these stimuli. The idea is to do something like AB or AXB analysis. Subjects try to determine which is a better, closer, or more correct usage of one item or which is more appropriate to use in a given context. Should be able to report on this soon.
No comments:
Post a Comment