PyCon HK 2026 Logo
PyCon HK 2026

AI 聽到聲音時,會聯想到乜嘢?

주목할 발표 · Jacky Chan

세션 소개

幾個月前,我問咗自己一個問題:如果我哋唔再叫 AI 模型「描述」音頻,而係問佢「聯想到啲乜嘢」,會發生咩事?

於是,我開始咗一個 Side Project,嘗試用 Raw Audio 直接 Probing 多模態 LLM(Multimodal LLM)。透過 Repeated Prompting 同 Position-Weighted Aggregation,去觀察模型會產生咩層次嘅聯想。

喺實驗過程中,我發現唔同模型對同一段聲音嘅反應有好大差異。呢個研究方法,其實好有潛力發展成一種全新嘅 Model Evaluation(模型評估)工具。

而家,我希望將呢個研究方向推向開源。未來幾個月我會推出一個開放嘅「聯想收集庫」,等大家可以上傳聲音、分享自己嘅聯想,從而對比人類同 AI 嘅思維分別,一齊建立一個好玩嘅開放數據集。

呢個 Talk 會分享我嘅實驗過程、主要發現,同埋對未來開源計劃嘅諗法。希望有興趣嘅朋友可以一齊嚟討論同貢獻!

발표자 소개

Jacky Chan

Jacky Chan

I am Jacky Chan, an researcher in Hong Kong who enjoys building and experimenting with local AI systems. I have been a regular participant in Hong Kong’s open source and PyCon community events.

My current work focuses on probing multimodal models to understand the associations they form with sound. I am also planning to open source parts of this work in the near future, with the goal of creating a community-driven audio association collection platform.

This is my first time proposing a talk at PyCon HK.

주요 연사로 돌아가기