------------
Thursday, June 27th
------------
I selected the audio files that I would use out of all of the stimuli that I collected. I then began writing phone and word-level textgrid annotations for them. The process was too tedious and taking too much time, so I used the penn forced aligner to generate the textgrids. Then all I needed to do was to fine-tune the textgrids. I finished most but not all of the textgrids by the end of the day.
------------
Friday, June 28th
------------
Not a very productive morning. Got an email from the IRB representative in the early afternoon that my ethics training was out-of-date. I did the ethics training which took a full 5 hours. After that I had headache and I called it a day.
------------
Sat - Sun
------------
Finished one or two outstanding modules in the ethics training.
I intended to work on the textgrid annotations but I ended up not having a productive weekend. Next weekend I need to go to the office.
------------
Monday, July 1st
------------
Finally finished up the textgrids. I went to generate the stimuli. No problem with the duration resynthesis but the pitch resynthesis crashed. The bug was nestled pretty deep but I had actually previously written a comment in the section where the bug originated.
The problem was that I was taking a constant number of pitch samples from every labeled region in the textgrid. This works fine as long as there are enough samples. However, the rest of the code assumes that there are a constant number of samples for each region. If there are more samples in the target audio than the source audio, no crash occurs. However, if there are more samples in the source audio than the target, then the code tries to map those last few samples to samples in the target audio that don't exist and the script crashes.
In either case, samples are mapped to non-corresponding samples, which is not what we want.
As a result, I changed it so that the pitch samples, and their time stamps, extracted from each region are stored in isolation until it is time to resynthesize, at which time they are pooled together (unlike the original code, where all samples were never separated out).
Additionally, in each region, for the source and target audio, I find the number of samples in both and pick whichever is smaller. I take this many samples from both. For the one with a larger number of samples, I sample at evenly spaced intervals such that the endpoints are preserved.
Before the pitch samples in the source and target audio are lined up, pitch values of 0 are removed.
After finishing with the code I finished for the day. However, the code still needs to be tested and run on the rest of the data set. I will pick a few random samples from both the duration and the pitch data and ensure that the interpolation is occurring as expected. Auditorily, it seems to be doing the right thing, but this needs to be verified sample-by-sample.
------------
Tuesday, July 2nd
------------
On account of rain I ended up staying at home and having a not-very-productive day. I need to go to the office every day, even if that means taking the bus on rainy days.
No comments:
Post a Comment