Publications
Export 28 results:
Author Title Type [ Year
] Filters: First Letter Of Last Name is R [Clear All Filters]
.
2026. Aligning Generative Music AI with Human Preferences: Methods and Challenges. Proceedings of AAAI, senior member track.
2511.15038v1.pdf (417.24 KB)
.
2026. Emerging AI Technologies for Music: Towards Controllable, Collaborative, and Creative Systems. Proceedings of Machine Learning Research, PMLR 303:1-5, 2026.
bhandari26a.pdf (161.47 KB)
.
2026. Generative AI in Education for SDG 4: Insights from Indonesia and Kazakhstan. Proceedings of the Pacific Asia Conference on Information Systems (PACIS)..
.
2026. KARMA-MV: A Benchmark for Causal Question Answering on Music Videos. arXiv:2605.08175.
2605.08175v1.pdf (3.32 MB)
.
2026. Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction. Conference on AI Music Creativity (AIMC).
.
2026. MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection. IEEE Tencon.
.
2026. nnAudio 2: Overcoming Dynamic Compilation Barriers and Transform Inconsistencies. Conference on AI Music Creativity (AIMC).
.
2026. SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering. Proceedings of ICML.
2508.03448v2.pdf (3.31 MB)
.
2026. Text2midi-InferAlign: Improving Symbolic Music Generation with Inference-Time Alignment. ICASSP.
2505.12669v1.pdf (360.69 KB)
.
2026. Text2Score: Generating Sheet Music From Textual Prompts. arXiv:2605.13431.
2605.13431v1.pdf (395.67 KB)
.
2026. Text2Score: Generating Sheet Music From Textual Prompts. arXiv:2605.13431.
2605.13431v1.pdf (395.67 KB)
.
2025. JamendoMaxCaps: A Large Scale Music-caption Dataset with Imputed Metadata. Proceedings of IJCNN, Rome, Italy.
.
2025. SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning. Proceedings of the 6th Conference on AI Music Creativity (AIMC 2025), Brussels, Belgium, September 10th - 12th, 2025.
.
2025. Text2midi: Generating Symbolic Music from Captions. Proceedings of AAAI, Philadelphia.
2412.16526v2.pdf (569.51 KB)
.
2024. MidiCaps — A large-scale MIDI dataset with text captions. ISMIR.
2406.02255v1.pdf (699.83 KB)
.
2024. MIRFLEX: Music Information Retrieval Feature Library for Extraction. ISMIR, Late Breaking Demos.
2411.00469v1.pdf (89.86 KB)
.
2022. EmoMV: Affective Music-Video Correspondence Learning Datasets for Classification and Retrieval. Information Fusion.
SSRN-id4189323.pdf (2.01 MB)
.
2022. HEAR 2021: Holistic Evaluation of Audio Representations. Proceedings of Machine Learning Research (PMLR): NeurIPS 2021 Competition Track.
2203.03022.pdf (406.58 KB)
.
2022. Predicting emotion from music videos: exploring the relative contribution of visual and auditory information to affective responses. Arxiv preprint.
.
2022. Single Image Video Prediction with Auto-Regressive GANs. Sensors. 22:3533.
.
2022. A white paper on cyberphysical learning. White paper, Singapore University of Technology and Design.
LSL_WhitePaper_Cyber-physical-Campus-Higher-Education.pdf (6.98 MB)
.
2021. AttendAffectNet – Emotion Prediction of Movie Viewers Using Multimodal Fusion with Self-attention. Sensors. Special issue on Intelligent Sensors: Sensor Based Multi-Modal Emotion Recognition.
sensors-21-08356.pdf (1.03 MB)
.
2021. AttendAffectNet: Self-Attention based Networks for Predicting Affective Responses from Movies. Proceedings of the International Conference on Pattern Recognition (ICPR2020).
2010.11188.pdf (7.07 MB)
.
2020. Regression-based music emotion prediction using triplet neural networks. Proceedings of the International Joint Conference on Neural Networks (IJCNN).
2001.09988.pdf (777.31 KB)
.
2019. Latent space representation for multi-target speaker detection and identification with a sparse dataset using Triplet neural networks. IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2019).
1910.01463.pdf (934.76 KB)
.
2019. Multimodal Deep Models for Predicting Affective Responses Evoked by Movies. The 2nd International Workshop on Computer Vision for Physiological Measurement as part of ICCV. Seoul, South Korea. 2019.
1909.06957.pdf (836.3 KB)
.
2018. Blacklisted speaker identification using triplet neural networks. MCE2018 competition.
SUTD_description.pdf (133.08 KB)