Publications

Export 160 results:
Author Title Type [ Year(Desc)]
2024
Melechovsky J., Mehrish A., Sisman B., Herremans D..  2024.  Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training. Proc. of IEEE Tencon, Singapore.
Melechovsky J., Mehrish A., Sisman B., Herremans D..  2024.  Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder. Proc. of IEEE Tencon, Singapore.
Lanzendörfer L.A., Lu T., Perraudin N., Herremans D., Wattenhofer R..  2024.  Coarse-to-Fine Text-to-Music Latent Diffusion. Audio Imagination: NeurIPS 2024 Workshop.
Melechovsky J., Mehrish A., Sisman B., Herremans D..  2024.  DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech. Audio Imagination: NeurIPS 2024 Workshop.
Ong J., Herremans D..  2024.  DeepUnifiedMom: Unified Time-series Momentum Portfolio Construction via Multi-Task Learning with Multi-Gate Mixture of Experts. arXiv:2406.08742. PDF icon 2406.08742v1.pdf (1.06 MB)
Wang K., Herremans D..  2024.  DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage. Proc. of IEEE Tencon, Singapore.
Chow D., Herremans D..  2024.  Gamification and skills tree. Trends and Foresight Report on Cyber-Physical Learning.
Melechovsky J., Roy A., Herremans D..  2024.  MidiCaps — A large-scale MIDI dataset with text captions. ISMIR. PDF icon 2406.02255v1.pdf (699.83 KB)
Chopra A., Roy A., Herremans D..  2024.  MIRFLEX: Music Information Retrieval Feature Library for Extraction. ISMIR, Late Breaking Demos. PDF icon 2411.00469v1.pdf (89.86 KB)
Ong J..  2024.  Modern Portfolio Construction with Advanced Deep Learning Models. SUTD. PhDPDF icon Joel_Ong_Thesis.pdf (3.44 MB)
Melechovsky J, Guo Z, Ghosal D, Majumder N, Herremans D, Poria S.  2024.  Mustango: Toward Controllable Text-to-Music Generation. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). pages 8293–8316. PDF icon 2311.08355 (1).pdf (11.38 MB)
Lam P., Zhang H., Chen N.F, Sisman B., Herremans D..  2024.  SNIPER Training: Variable Sparsity Rate Training For Text-To-Speech. Proc. of IEEE Tencon, Singapore. PDF icon 2211.07283.pdf (435.22 KB)
Kang J, Poria S, Herremans D..  2024.  Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model. Expert Systems with Applications. PDF icon 2311.00968.pdf (5.51 MB)
2025
Melechovsky J..  2025.  Analysis and Synthesis of Audio with AI: from Neurological Disease to Accented Speech and Music. PDF icon thesis_Jan.pdf (26.4 MB)
Kang J., Herremans D..  2025.  Are we there yet? A brief survey of Music Emotion Prediction Datasets, Models and Outstanding Challenges IEEE Transactions on Affective Computing. PDF icon 2406.08809v1.pdf (156.19 KB)
Luo J., Yang X., Herremans D..  2025.  BandCondiNet: Parallel Transformers-based Conditional Popular Music Generation with Multi-View Features. Expert Systems with Applications. 130059PDF icon 2407.10462v2.pdf (2.6 MB)
Lanzendörfer L.A., Lu T., Perraudin N., Herremans D., Wattenhofer R..  2025.  Coarse-to-Fine Text-to-Music Latent Diffusion. Proceedings of ICASSP.
Song M., Liu R., Wang X, Jiang Y, Xie P, Huang F, Zhou J, Herremans D., Poria S..  2025.  Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics. arXiv:2510.05137.
Tripathi A., Patle V., Jain A., Pundir A., Menon S., A. Singh K, Herremans D..  2025.  End-to-End Text-to-SQL with Dataset Selection: Leveraging LLMs for Adaptive Query Generation. Proceedings of IJCNN, Rome, Italy.
Guo R, Herremans D..  2025.  An exploration of controllability in symbolic music infilling. IEEE Access.
Herremans D., Low K.W..  2025.  Forecasting Bitcoin Volatility Spikes from Whale Transactions and Cryptoquant Data Using Synthesizer Transformer Models. IEEE Access. 13:117788-117807.PDF icon SSRN-id4247684.pdf (5.05 MB)
Tripathi A., A. Singh K, Surya R., Gupta A., Veikho S.L., Herremans D., Bisane S..  2025.  HHNAS-AM: Hierarchical Hybrid Neural Architecture Search using Adaptive Mutation Policies. arXiv:2508.14946.
Bhandari K., Chang S., Lu T., Enus F.R, Bradshaw L.B, Herremans D., Colton S..  2025.  ImprovNet: Generating Controllable Musical Improvisations with Iterative Corruption Refinement. Proceedings of IJCNN.
Roy A., Liu R., Lu T., Herremans D..  2025.  JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata. Proceedings of IJCNN, Rome, Italy.
Le D-V-T, Bigo L., Keller M., Herremans D..  2025.  Natural Language Processing Methods for Symbolic Music Generation and Information Retrieval: a Survey. ACM Computing Surveys. PDF icon 2402.17467.pdf (1.01 MB)
Lam P., Zhang H., Chen N.F, Sisman B., Herremans D..  2025.  PRESENT: Zero-Shot Text-to-Prosody Control. IEEE Signal Processing Letters. PDF icon 2408.06827v1.pdf (367.55 KB)
Wei M., Modrzejewski M., Sivaraman A., Herremans D..  2025.  Prevailing Research Areas for Music AI in the Era of Foundation Models.
Herremans D..  2025.   Royalties in the age of AI: paying artists for AI-generated songs. WIPO Magazine.
Wickramasinghe S., Das B., Herremans D..  2025.  Smart Timing for Mining: A Deep Learning Framework for Bitcoin Hardware ROI Prediction. PDF icon 2512.05402v1.pdf (908 KB)
Chopra A., Roy A., Herremans D..  2025.  SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning. Proceedings of the 6th Conference on AI Music Creativity (AIMC 2025), Brussels, Belgium, September 10th - 12th, 2025.
Bhandari K., Roy A., Wang K., Puri G., Colton S., Herremans D..  2025.  Text2midi: Generating Symbolic Music from Captions. Proceedings of AAAI, Philadelphia. PDF icon 2412.16526v2.pdf (569.51 KB)
Sockalingam N., Lo K., Teo J., Wei C.C., Chow D., Herremans D., Jun M.L.M., Kurniawan O., Wang Y., Leong P.K.  2025.  Towards the future of education: cyber-physical learning. Discover Education. 4:1–16.
Kang J., Herremans D..  2025.  Towards Unified Music Emotion Recognition across Dimensional and Categorical Models.
2026
Herremans D., Roy A..  2026.  Aligning Generative Music AI with Human Preferences: Methods and Challenges. Proceedings of AAAI, senior member track. PDF icon 2511.15038v1.pdf (417.24 KB)
Husain J.A., Herremans D..  2026.  APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music. arXiv:2605.03395. PDF icon 2605.03395v1 (1).pdf (292.45 KB)
Melechovsky J., Novotny M., Tykalova T., Klempir J., Herremans D..  2026.  Development of Interpretable Deep Learning-based Segmentation Algorithm for Automated Assessment of Oral Diadochokinesis in Progressive Neurological Diseases. Journal of Speech, Language, and Hearing Research.
Puri G., Socklingam N., Herremans D..  2026.  Digital Lifelong Learning in the Age of AI: Trends and Insights .
Bhandari K., Roy A., Colton S., Herremans D..  2026.  Emerging AI Technologies for Music: Towards Controllable, Collaborative, and Creative Systems. Proceedings of Machine Learning Research, PMLR 303:1-5, 2026. PDF icon bhandari26a.pdf (161.47 KB)
Kadir N., Herremans D..  2026.  A Functional Taxonomy of Intelligent Systems in Education.
A. Putri M, Saide S., D. Riau K, Herremans D..  2026.  Generative AI in Education for SDG 4: Insights from Indonesia and Kazakhstan. Proceedings of the Pacific Asia Conference on Information Systems (PACIS)..
Ghosh A., Roy A., Herremans D..  2026.  KARMA-MV: A Benchmark for Causal Question Answering on Music Videos. arXiv:2605.08175. PDF icon 2605.08175v1.pdf (3.32 MB)
Liu R., Roy A., Herremans D..  2026.  Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction. Conference on AI Music Creativity (AIMC).
Song M., Pala T.D, Jin W., Zadeh A., Li C., Herremans D., Poria S..  2026.  Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social Interactions. Proceedings of ICLR.
Lu T., Geist C-M, Melechovsky J., Roy A., Herremans D..  2026.  MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection. IEEE Tencon.
Roy A, Liang J, Herremans D.  2026.  nnAudio 2: Overcoming Dynamic Compilation Barriers and Transform Inconsistencies. Conference on AI Music Creativity (AIMC).
Jiang Z., Yeo S., Herremans D., Perrault S..  2026.  Scaffolded Vulnerability: Chatbot-Mediated Reciprocal Self-Disclosure and Need-Supportive Interaction in Couples. Proceedings of CHI.
Melechovsky J., Mehrish A., Roy A., Herremans D..  2026.  SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering. Proceedings of ICML. PDF icon 2508.03448v2.pdf (3.31 MB)
Roy A., Puri G., Herremans D..  2026.  Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment. ICASSP. PDF icon 2505.12669v1.pdf (360.69 KB)
Bhandari K., Chang S., Roy A., Ronchini F., Benetos E., Herremans D., Colton S..  2026.  Text2Score: Generating Sheet Music From Textual Prompts. arXiv:2605.13431. PDF icon 2605.13431v1.pdf (395.67 KB)
Wang K, Quek B-K, Goh J, Herremans D.  2026.  To Embody or Not: The Effect Of Embodiment On User Perception Of LLM-based Conversational Agents. IEEE Tencon.

Pages