A METHOD FOR EVALUATING THE CONFIDENCE OF MUSIC GENERATION BASED ON BAYESIAN TRANSFORMER. 168-181 SI

Yuan Li

References

  1. [1] R. Ogden, J. Wearden, L. Jones, L. B. Silva, M. Phillips,and J. O. Martins, The influence of tonality, tempo, andmusical sophistication on the listener’s time-duration estimates,Quarterly Journal of Experimental Psychology, 77(9), 2024,1846–1864.
  2. [2] G. Carraturo, V. Pando-Naude, M. Costa, P. Vuust, L. Bonetti,and E. Brattico, The major-minor mode dichotomy in musicperception, Physics of Life Reviews, 52(3), 2025, 80–106.
  3. [3] R. Hake, M. B¨urgel, N. K. Nguyen, A. Greasley, D. M¨ullensiefen,and K. Siedenburg, Development of an adaptive test ofmusical scene analysis abilities for normal-hearing and hearing-impaired listeners, Behavior Research Methods, 56(6), 2024,5456–5481.
  4. [4] H. Abudukelimu, J. Chen, Y. Liang, A. Abulizi, and A.Yasen, SymforNet: Application of cross-modal informationcorrespondences based on self-supervision in symbolic musicgeneration, Applied Intelligence, 54(5), 2024, 4140–4152.
  5. [5] J. M. Roche, S. D. Morgan, and S. Fisk, Gender stereo-types drive perceptual differences of vocal confidence, TheJournal of the Acoustical Society of America, 151(5), 2022,3031–3042.
  6. [6] A. Ashurov, Z. Yi, and Y. M. Li, Concatenation-basedpre-trained convolutional neural networks using attentionmechanism for environmental sound classification, AppliedAcoustics, 216, 2024, 109759.1–109759.12.
  7. [7] R. Ogden, J. Wearden, L. Jones, M. N. Plastira, M. P.Michaelides, and M. N. Avraamides, Music and speech timeperception of musically trained individuals: The effects of audiotype, duration of musical training, and rhythm perception,Quarterly Journal of Experimental Psychology, 77(9), 2024,1835–1845.
  8. [8] H. Park, Y. Chung, and J. H. Kim, Deep neural networks-basedclassification methodologies of speech, audio and music, andits integration for audio metadata tagging, Journal of WebEngineering, 22(1), 2023, 1–25.
  9. [9] E. V. Thomas, Use of the bias-corrected parametric bootstrapin sensitivity testing/analysis to construct confidence boundswith accurate levels of coverage, Journal of Quality Technology,55(4), 2023, 22.
  10. [10] J. D. V. Quiros, L. Cabrera-Quiros, C. Oertel, and H. Hung,Impact of annotation modality on label quality and modelperformance in the automatic assessment of laughter in-the-wild, IEEE Transactions on Affective Computing, 15(2), 2024,519–534.
  11. [11] A. Serrurier, C. Neuschaefer-Rube, and R. R¨ohrig,. Past andtrends in cough sound acquisition, automatic detection andautomatic classification: A comparative review, Sensors (Basel,Switzerland), 22(8), 2022, 1–30.
  12. [12] Z. Guo, S. Wu, M. Ohno, and R. Yoshida, Bayesian algorithm forretrosynthesis, Journal of Chemical Information and Modeling,60(10), 2020, 4474–4486.
  13. [13] Y. C. Tsai and F. C. Lin, Paraphrase generation modelintegrating Transformer architecture, part-of-speech features,and pointer generator network, IEEE Access, 11(3), 2023,30109–30117.
  14. [14] T. Colafiglio, C. Ardito, P. Sorino, D. Lof`u, F. Festa, T. D.Noia, and E. Di Sciascio, NeuralPMG: A neural polyphonicmusic generation system based on machine learning algorithms,Cognitive Computation, 16(5), 2024, 2779–2802.
  15. [15] K. Garba, T. Kolajo, and J. B. Agbogun, A Transformer-basedapproach to nigerian pidgin text generation, InternationalJournal of Speech Technology, 27(4), 2024, 1027–1037.
  16. [16] N. Yadav, A. Kumar Singh, and S. Pal, Improved self-attentive musical instrument digital interface content-basedmusic recommendation system, Computational Intelligence,38(4), 2022, 1232–1257.
  17. [17] A. Diaz-Arias and D. Shin, ConvFormer: Parameter reductionin Transformer models for 3D human pose estimation byleveraging dynamic multi-headed convolutional attention, TheVisual Computer, 40(4), 2024, 2555–2569.
  18. [18] D. Tiwari and B. Nagpal, KEAHT: A knowledge-enrichedattention-based hybrid Transformer model for social sentimentanalysis, New Generation Computing, 40(4), 2022, 1165–1202.
  19. [19] O. Mahmoudi, M. Filali-Bouami, and M. Benchat, Speechrecognition based on the Transformer’s multi-head attentionin arabic, International Journal of Speech Technology, 27(1),2024, 211–223.
  20. [20] S. Ghaith, Deep context Transformer: Bridging efficiencyand contextual understanding of Transformer models, AppliedIntelligence, 54(19), 2024, 8902–8923.
  21. [21] M. Kang, J. Park, H. Shin, J. Shin, and L. S.Kim, ToEx:Accelerating generation stage of Transformer-based languagemodels via token-adaptive early exit, IEEE Transactions onComputers, 73(9), 2024, 2248–2261.
  22. [22] M. Mukhopadhyay and S. Bhattacharya, Bayes factorasymptotics for variable selection in the Gaussian processframework, Annals of the Institute of Statistical Mathematics,74(3), 2022, 581–613.
  23. [23] D. Ming, D. Williamson, and S. Guillas, Deep Gaussian processemulation using stochastic imputation, Technometrics, 65(2),2023, 150–161.
  24. [24] A. Ozerov, P. Philippe, F. Bimbot, and R. Gribonval,Adaptation of Bayesian models for single-channel sourceseparation and its application to voice/music separation inpopular songs, IEEE Transactions on Audio, Speech, andLanguage Processing, 15(5), 2007, 1564–1578.

Important Links:

Go Back