Auditorium
Conference Schedule
All times: Hong Kong (UTC+8)
No sessions match these filters.
10:00
10:10
Auditorium
Group photo
10:15
Auditorium
Speakers: Paul Everitt
10:45
Auditorium
Stage changeover
11:20
Auditorium
Panel setup
LT1 (AI, Data)
Move to the small rooms
LT2 (Software)
Move to the small rooms
LT3 (Community)
Move to the small rooms
11:25
Auditorium
Speakers: Calvin Tsang, Soojin Yoon (Luna)
LT2 (Software)
Speakers: Tristan McKinnon
LT3 (Community)
Speakers: Yuichiro Tachibana
11:55
LT1 (AI, Data)
Break
LT2 (Software)
Break
LT3 (Community)
Break
12:00
LT1 (AI, Data)
Speakers: Su Myat Noe
LT3 (Community)
Speakers: Hon Kwan Shun Quinson
12:30
LT1 (AI, Data)
Break
LT2 (Software)
Get ready for the short talks
LT3 (Community)
Break
12:35
LT1 (AI, Data)
Speakers: Jacky Chan
LT2 (Software)
Speakers: Shardul Deshpande
LT3 (Community)
Speakers: jungmir
12:50
LT1 (AI, Data)
Lunch
LT2 (Software)
Lunch
LT3 (Community)
Lunch
13:50
LT1 (AI, Data)
Speakers: Tarun Jain
LT3 (Community)
Speakers: Nizar Akbar Meilani
14:20
LT1 (AI, Data)
Move between sessions
LT2 (Software)
Break
LT3 (Community)
Break
14:25
LT1 (AI, Data)
Speakers: Alan, Ka Hei Ng, William Leung
LT2 (Software)
Speakers: Dr. Adrian Tam
LT3 (Community)
Speakers: satyam soni
14:55
LT1 (AI, Data)
Move between sessions
LT2 (Software)
Break
LT3 (Community)
Break
15:00
LT1 (AI, Data)
Speakers: Indy Ho, Ken Lee
LT2 (Software)
Speakers: Wenxin Jiang
15:30
LT1 (AI, Data)
Tea break / Posters (display area TBC)
LT2 (Software)
Tea break / Posters (display area TBC)
LT3 (Community)
Tea break / Posters (display area TBC)
16:10
LT1 (AI, Data)
Speakers: Kazelegendo
LT3 (Community)
Speakers: Soojin Yoon (Luna)
16:40
LT1 (AI, Data)
Get ready for shared closing (room TBC)
LT2 (Software)
Join shared closing: lightning talks + lucky draw (room TBC)
LT3 (Community)
Join shared closing: lightning talks + lucky draw (room TBC)
16:45
LT1 (AI, Data)
Speakers: Clarissa Gunawan
16:50
LT1 (AI, Data)
Lightning talks (shared session; room TBC)
17:15
LT1 (AI, Data)
Closing and Sunday information (room TBC)
Speakers
Paul is a Developer Advocate at JetBrains, focusing on Python and the Web. Before that, Paul was a co-founder of Zope Corporation, taking the first open source application server through $14M of funding. Paul has bootstrapped both the Python Software Foundation and the Plone Foundation. Prior to that, Paul was an officer in the US Navy, starting www.navy.mil in 1993.
About this session
Python is in a moment. Our entire profession is in a moment. But Python has more power than we think we do: us. In this talk we cover how "us" came to be, what the age of agentic coding has done to our profession and community, and close with what we can do about it: a Python AI project from us, for us.
This is intended to be a keynote. I want it to be a celebration, a moment of group catharsis, and a rallying cry. The first third will come from my standard "Python 1994" keynote. The middle third from my April agentic engineering talk at Andrew Ng's AI Dev conference. The last third comes from a conversation and plan involving a number of key Python people this year.
Speakers
Pyladies Hong Kong is established on 2025 There are including active members: Cathy, Jenny, Cintia and Daisy . , PyLadies Seoul: Luna, Pyladies Tokyo: Maaya
- PyLadies Seoul Organizer
- Django Korea Organizer
- AWS Women in Cloud Organizer
- Software engineer at SocraAI (formerly Riiid), an AI edtech company
- Pythonista using Python and Django
- Experienced speaker at PyCon
- Proud owner of a cute dog
- Korean, mainly communicating in English, and learning a little Japanese.
About this session
PyLadies communities across Asia are creating welcoming spaces for learning Python, developing technical skills, and empowering future technology leaders. In this panel discussion, organizers from PyLadies Seoul, PyLadies Tokyo, and PyLadies Hong Kong will share how they build and sustain their communities, organize impactful events, overcome common challenges, and foster meaningful connections among members. Through stories, lessons learned, and practical experiences, attendees will gain insights into community leadership, volunteer engagement, and the growing role of PyLadies in strengthening the Python ecosystem across Asia. Whether you are a community organizer, aspiring volunteer, conference speaker, or Python enthusiast, this session offers valuable perspectives on building inclusive and sustainable tech communities.
30-Minute Agenda Time Topic 2 min Moderator Introduction 9 min Community Spotlight (3 min each chapter) 10 min Guided Panel Discussion 7 min Audience Q&A 2 min Closing & Call for CollaborationCommunity Spotlight (3 min each)
Each chapter answers: Who are you? How large is your community? What activities do you run? One achievement you're proud of.
Guided Discussion Topics (10 min) Topic 1 – Growing a Community How do you attract new members? What event formats work best? Topic 2 – Challenges Volunteer recruitment Organizer sustainability Finding speakers Retaining members Topic 3 – Culture & Collaboration What's unique about your local community? What can PyLadies chapters learn from each other?
Discussion: Why people join PyLadies Success stories from members Becoming first-time speakers Leadership development Open source participation Building confidence in tech
Speakers
I build data systems for environments where a single compromised credential can expose millions of patient records.
By day I'm a Lead Healthcare Data Engineer at Axle Informatics, designing zero-trust pipelines for genomic and clinical data at NIH scale. I also run Deterministic Systems Lab, where I do independent security research, and I'm a doctoral student in AI/ML at George Washington University.
Most of my work orbits a core problem: identity and trust in high-velocity data systems. I developed the Identity-Per-Transaction protocol: credentials that are cryptographically scoped to a single transaction and gone the moment it completes, published in IEEE BigDataSecurity. From there I've been building outward: cryptographic lineage so every transformation is auditable, and purpose-aware query governance using graph neural networks to enforce data use restrictions before a result ever leaves a clinical knowledge graph.
I spend a lot of time thinking about what 'secure by design' actually means when your data is regulated, your pipelines are serverless, and your threat model includes people who work there.
Happy to talk shop on any of it; find me after the talk or just grab me in the hallway.
About this session
Most "zero trust" data pipelines are perimeter security with a new label: authenticate once, trust the process for hours, and one leaked credential exposes the whole bucket. This talk shows a different model: issue a narrowly scoped, 900-second credential per transaction instead of a standing role per user. I build it in plain Python (boto3, STS, a Lambda trigger), using a healthcare pipeline under FedRAMP High as the worked example, and I'm straight about where it holds and where it doesn't. You'll leave knowing how to scope a credential to a single S3 object and why that collapses the blast radius of a breach by over 99%.
The default identity model in cloud data engineering is Identity-Per-User: a long-lived role attached to a persistent cluster. It's convenient, and it's a liability. The moment the process boots, least privilege is already gone, and a single leaked credential exposes everything the role can touch for the life of the session.
This talk walks through the alternative I formalized as Identity-Per-Transaction. Instead of a standing role, an identity broker synthesizes an IAM policy at trigger time, scoped to the exact object that fired the event, and hands back an STS credential that expires in 900 seconds. The compute is ephemeral; the permission is disposable.
I build it in plain Python and structure the talk around a three-zone "clean room": a dirty landing zone with no read access, an airlock with no internet gateway where ephemeral functions run, and a clean zone that only ever receives tokenized data. The code that matters fits on two slides: how the broker builds a per-object policy inline, and how the scoped session is assumed.
Then the honest part. I'm specific about what we measured (blast radius down over 99.9%, sub-second issuance overhead, idle cost to zero) and clear about the limits: this is a pilot awaiting a penetration test, validated on batch workloads rather than at streaming scale, and the wrong tool for tight low-latency loops.
Speakers
Yuichiro is a professional software developer with a deep passion for open source software. After working on software development for several machine learning startups, he joined Hugging Face in 2023. In 2026, he joined the Research Center for Advanced Science and Technology, The University of Tokyo.
He develops and maintains several OSS projects, including Streamlit-WebRTC, Stlite, and Awesome Emacs Keymap, and contributes to other OSS projects such as Streamlit and Gradio. He has also been actively involved in the global PyCon community, attending PyCons around the world and sharing his work through talks at several Python conferences
About this session
Browser ports of Python web frameworks need more than Python running in WebAssembly through Pyodide. For example, Stlite - an in-browser version of Streamlit - needs a runtime layer that maps HTTP requests, WebSocket messages, and other browser-side events into Python code running on Pyodide. From reading Shinylive's ASGI bridge and implementing similar layers for Stlite and Gradio-Lite, I became convinced that ASGI is a useful abstraction for this server-like layer. In this talk, I will explain how an ASGI bridge layer running in the browser can receive HTTP requests from JavaScript, return responses, handle WebSocket messages, drive lifespan events, and preserve PEP 567 context variables. I will also show how this applies to ASGI-based frameworks such as Starlette and FastAPI.
ASGI is usually discussed as the interface between Python web applications and servers such as Uvicorn or Hypercorn. Pyodide is usually discussed as a way to run Python inside the browser. This talk combines these two ideas: what happens if the “server” side of ASGI is implemented as a bridge layer inside a browser tab?
The talk is based on my experience creating Stlite, an in-browser version of Streamlit, and Gradio-Lite, an in-browser version of Gradio. In these projects, Python code runs on Pyodide, but the browser still needs a runtime layer that bridges HTTP requests, WebSocket messages, and other browser-side events to Python web framework code. After reading the relevant code in Shinylive and implementing similar layers myself in Stlite and Gradio-Lite, I learned that ASGI is a useful abstraction for this layer that reproduces the role of a web server in the browser.
I will first show a small Starlette/FastAPI-style example running in the browser. Then I will explain the bridge layer behind it: how HTTP requests are converted into ASGI scopes, how receive and send are connected to JavaScript, how WebSocket sessions are represented, how lifespan events are driven, and why PEP 567 context variables matter when multiple application instances share the same Pyodide runtime.
This is not only a deployment trick. It is also a way to understand ASGI more concretely. By seeing what has to be implemented on the “server” side, attendees can better understand what ASGI frameworks expect from their runtime. I will close with practical use cases such as static-hosted demos, runnable documentation, educational examples, and privacy-preserving browser apps, along with the limits of this approach.
Relevant projects and references:
- Stlite: https://github.com/whitphx/stlite
- Gradio-Lite implementation: https://github.com/gradio-app/gradio/pull/4402
- Pyodide: https://pyodide.org/
- ASGI specification: https://asgi.readthedocs.io/
Speakers
Su Myat Noe is a Project Researcher at the National Institute of Informatics (NII) in Tokyo, working on AI evaluation. Most of her days are spent in Python — building, breaking, and rebuilding evaluation pipelines for large language and vision-language models. She holds a PhD in Computer Science from the University of Miyazaki (2025) and came to AI through computer vision research on livestock tracking. Originally from Myanmar and based in Japan since 2019, she is a Board Member of Women in AI Myanmar and volunteers with GDG Tokyo. This is her first PyCon JP.
About this session
LLM-as-Judge has become the default way to evaluate AI systems — for safety, helpfulness, quality. It's also one of the easiest pieces of code to silently get wrong, producing plausible numbers that don't mean what you think. This talk tours failure modes I've hit while building LLM-as-Judge pipelines in Python: prompts that drift from your rubric, ignored reference answers, brittle score parsing, English-default prompts that quietly degrade on other languages. I'll walk through a minimal open-source judge pipeline, deliberately introduce each failure mode, and show the cheap sanity checks that catch them: perturbation tests, reference comparisons, distribution diagnostics. If you trust an LLM to grade your model's outputs, you should know how it can lie.
Copy-paste ready. Pretalx supports Markdown so the formatting will render on the public talk page:
LLM-as-Judge has quietly become the default way to evaluate AI systems — for safety, helpfulness, quality, almost everything. It's also one of the easiest pieces of code to silently get wrong. Your pipeline runs, your numbers look reasonable, your correlations are plausible. Nothing crashes. And the scores mean something entirely different from what you think they mean. This talk is a Python field guide to the failure modes I've collected from building (and breaking) LLM-as-Judge pipelines across multiple research projects over the past year. I'll walk through a minimal open-source judge pipeline, then deliberately introduce each failure mode and show the cheap sanity check that catches it before it ships. What you'll see:
Prompt drift — when your few-shot examples teach the model to grade a different rubric than the one you wrote. Reference-answer blindness — when the model ignores the ground truth you carefully constructed and scores on its own opinion. Score format brittleness — when half your scores silently become None, or "10" gets parsed as "1". Cross-language assumptions — when an English judge prompt quietly degrades on non-English inputs.
For each failure mode I'll show: a minimal buggy version, the detection check that exposes it, and the fix. Every example runs from a public GitHub repo I'll build before the conference — no GPU needed, no API spend beyond a few dollars. Attendees can fork it, run the bugs themselves, and contribute their own failure modes. Who this is for: anyone using (or considering) LLM-as-Judge in their Python workflow. You don't need experience with safety evaluation or vision-language models — the failure modes are general. You'll leave with a small toolkit of sanity checks you can drop into your own evaluation pipeline tomorrow. What you'll take away:
A minimal reference implementation of an LLM-as-Judge pipeline in Python. Four failure modes, each with buggy code, a detection check, and a fix. Five sanity-check patterns — perturbation, reference-swap, distribution diagnostics, parse-failure logging, and cross-language matched pairs — that you can drop into your own pipeline. A clearer sense of when LLM-as-Judge is the right tool, and when to reach for human eval, rule-based metrics, or programmatic checkers instead.
Speakers
Hon Kwan Shun Quinson is an M.Phil. student in Computer Science at The Chinese University of Hong Kong. His academic pursuits and research focus on the application of LLMs in software development, leveraging existing software tools to enhance the capability of AI agents in software development and general workflows.
About this session
When writing code, it is often good practice to write tests to ensure that the code is correct. Typically, these tests check the code one input at a time to see if the output is correct for each case. However, in practice, we want the code to work on a wide range of inputs, and it is sometimes impossible to test all cases. To mitigate this, property-based testing can be used to generate random but structured inputs to pass into the code, and the code can be tested on a much larger range of inputs to give developers better confidence that their code works. This talk will explain in more detail what property-based testing is, how to conduct property-based testing with the Hypothesis library, when it is useful and how it can be integrated into standard development workflows. Attendees will learn various property-based testing concepts and techniques to ensure that their inputs work on a wider range of scenarios.
Speakers
I am Jacky Chan, an researcher in Hong Kong who enjoys building and experimenting with local AI systems. I have been a regular participant in Hong Kong’s open source and PyCon community events.
My current work focuses on probing multimodal models to understand the associations they form with sound. I am also planning to open source parts of this work in the near future, with the goal of creating a community-driven audio association collection platform.
This is my first time proposing a talk at PyCon HK.
About this session
幾個月前,我問咗自己一個問題:如果我哋唔再叫 AI 模型「描述」音頻,而係問佢「聯想到啲乜嘢」,會發生咩事?
於是,我開始咗一個 Side Project,嘗試用 Raw Audio 直接 Probing 多模態 LLM(Multimodal LLM)。透過 Repeated Prompting 同 Position-Weighted Aggregation,去觀察模型會產生咩層次嘅聯想。
喺實驗過程中,我發現唔同模型對同一段聲音嘅反應有好大差異。呢個研究方法,其實好有潛力發展成一種全新嘅 Model Evaluation(模型評估)工具。
而家,我希望將呢個研究方向推向開源。未來幾個月我會推出一個開放嘅「聯想收集庫」,等大家可以上傳聲音、分享自己嘅聯想,從而對比人類同 AI 嘅思維分別,一齊建立一個好玩嘅開放數據集。
呢個 Talk 會分享我嘅實驗過程、主要發現,同埋對未來開源計劃嘅諗法。希望有興趣嘅朋友可以一齊嚟討論同貢獻!
Speakers
Shardul Deshpande works at Canonical on the Observability Stack (COS), building practical demos around observability, Kubernetes, and self-hosted AI infrastructure. This talk is based on a reproducible Multipass / MicroK8s / COS experiment that turns LLM behaviour into metrics, logs, traces, dashboards, alerts, and quality signals.
About this session
You put a FastAPI service in front of a local LLM, it returns 200 OK, and the model answers — but is it actually good? The first token might arrive seconds late, a single prompt might burn your token budget, or a reply might come back truncated, all while your HTTP and pod metrics stay green.
This talk shows how to instrument a Python LLM gateway with the OpenTelemetry GenAI semantic conventions — emitting spans, metrics, and trace-correlated logs straight from async streaming code — so you can actually see time-to-first-token, token burn, errors, and a simple quality signal. Then we run two local models on the same CPU and let the telemetry pick the winner, instead of guessing.
The problem
Most "run an LLM locally" tutorials stop at the first 200 OK. But a serving path can look perfectly healthy at the HTTP and container level while users wait too long for the first token, requests burn far more tokens than budgeted, responses come back truncated or malformed, or the wrong model quietly handles the request. Ordinary web metrics can't see any of it.
What we build (in Python)
- A FastAPI OpenAI-compatible gateway in front of Ollama, serving small models on CPU.
asyncupstreaming with httpx, including server-sent-events (SSE) streaming and a second code path that adapts Ollama's native API — so reasoning models can disable "thinking" while clients keep calling/v1/chat/completions.- Hand-rolled instrumentation with
opentelemetry-python: a tracer plus histograms and a counter, custom histogram buckets via SDKViews, trace-based exemplars (metrics recorded inside the active span so a Grafana spike links to the exact trace), and JSON logs that carrytrace_id/span_id.
The GenAI semantic conventions
OpenTelemetry graduated in CNCF in May 2026, but its GenAI semantic conventions are still marked Development. The talk is honest about that: the gateway opts in explicitly with OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental and emits attributes like gen_ai.request.model, gen_ai.usage.*, and finish reasons, rather than pretending the names are stable.
The demo
A small asyncio load generator replays recorded scenarios — normal, slow, expensive, error, and low-quality traffic — so every signal has something to show. On the dashboard you watch time-to-first-token, token burn, finish reasons, error rate, and an opt-in heuristic quality score move, while the platform never notices. We finish with a matched comparison of llama3.2:1b vs qwen3.5:2b on the same CPU and prompt, changing only the model — turning "which model should we run?" into a measured question.
Prompt/response capture is off by default; when enabled, a small redaction pass strips PII before anything is exported.
What you'll leave with
- A mental model for what to measure on an LLM call beyond status codes and latency.
- A concrete pattern for instrumenting
asyncFastAPI code withopentelemetry-pythonand the GenAI conventions. - A reproducible, CPU-only, fully open-source scaffold you can run on a laptop — no GPU, no vendor SaaS.
Audience & level
Intermediate Python developers. Helpful: basic FastAPI/async familiarity. Not required: Kubernetes — the backend (Prometheus/Loki/Tempo/Grafana via Canonical's COS) is shown as the payoff, and the talk introduces the observability concepts before the demo.
Outline (30 min)
- "The model answered" is not enough — 3 min
- The FastAPI gateway and its two async streaming paths — 6 min
- Instrumenting with
opentelemetry-python: spans, metrics, exemplars, correlated logs — 8 min - The GenAI semantic conventions, and what's still experimental — 3 min
- Live signals across the scenarios; one trace explaining a slow request — 6 min
llama3.2:1bvsqwen3.5:2b, decided by telemetry — 3 min- Limitations & wrap-up — 1 min
Speakers
Hello, I am Mir Jung, a Backend Engineer. I currently work at AhnLab, a security software company in South Korea, and I also serve as an organizer for PyCon Korea.
I love Python and am deeply interested in contributing to the community ecosystem. Recently, through my contributions to translating the official Python documentation, I have been thinking deeply about how technical knowledge can be accurately delivered without language barriers.
Do Anything with Python!
About this session
In the era of AI, machine translation makes any technical document seem effortless to read. If you ask an AI, it will even hand you sample code, making it seem completely unnecessary to read the documentation at all. Does this mean it is no longer meaningful for developers to invest their precious time manually translating the official Python documentation?
This session begins with the candid experience of a developer who, struggling with English, relied on AI to attempt translation contributions. Through a vivid real-world example (e.g., the trap of the unless clause), we will explore how seemingly flawless but technically out-of-context machine translations can completely distort the logic of the code. Rather than helping, these disjointed documents often become critical bottlenecks that hinder the growth of newcomers to the Python ecosystem.
Moving beyond simple 1:1 word replacement, we will share the journey of reconstructing machine language into true "engineering language" through rigorous community PR reviews. This highlights that contributing to translations is not mere text conversion, but a vital software engineering task that fortifies the Python ecosystem. We invite anyone wishing to overcome language barriers and the pitfalls of machine translation to take a meaningful first step into the Python community.
Speakers
Tarun Jain is a Founding Engineer at Kaivid Labs, Google Developer Expert in AI, and Qdrant Distinguished Ambassador. Tarun has contributed to Google Summer of Code 2024 at Red Hen Lab and Google Summer of Code 2023 at caMicroscope. He is a content creator on YouTube with the channel name: AI with Tarun.
About this session
Pandas works until you hit memory limits, slow joins, or null coercion bugs. Polars fixes most of these by design: columnar memory layout, strict types, lazy evaluation, and real parallelism.
We will start with familiar Pandas style analytics tasks, then showcase selected parts using Polars to show where lazy execution, query optimization, and memory-efficient execution become useful. The focus is not on replacing Pandas everywhere, but on recognizing the point where Polars gives clearer, faster, or more maintainable workflows.
From there, we look at sandboxed execution: running Polars in an isolated environment to safely execute dynamically generated transformation code that can also be used as a tool for those building in Agents for Data Analytics and Visualization Agentic workflow.
Many Python data workflows begin with Pandas because it is familiar, flexible, and deeply integrated into the ecosystem. That is still a reasonable default. However, as datasets grow, transformations become more complex, and pipelines move closer to production, the trade-offs become more apparent.
This talk uses practical examples to compare Pandas and Polars from the perspective of everyday analytics work. We will cover where Pandas remains a good choice, where Polars starts to make sense, and how to think about the migration path without rewriting everything at once. The main focus is on executing Polars schema within the sandbox, such as a REPL-based tool when the dataset is huge and latency matters.
Do note: This talk explains why Polars is worth understanding, not as a replacement for every Pandas use case, but as a practical tool for cases where query planning, lazy execution, and memory-efficient processing matter.
Speakers
Nizar Akbar Meilani is a systems engineer with deep expertise in Linux infrastructure, DevOps, SRE, and Infrastructure as Code. His analytical approach to systems problems—rooted in a passion for precise data over assumptions—was forged through years of reverse engineering game binaries and analyzing complex software behavior at the binary level.
Nizar has presented at PyCon APAC 2024 ("Enhancing Actively Attacked WordPress Vulnerability Detection with Python"), PyCon Hong Kong 2025 ("IPList to BPFRule: A Python DDoS Mitigation Framework for Domain-Level Attacks via XDP_HOOK of bpfilter"), and PyCon Indonesia 2025 ("Python meets bpfilter and safeline: Low Cost HTTP Flood Filtering"), sharing insights on Python-powered infrastructure security and network filtering. His current focus is expanding from kernel-level traffic filtering into the broader eBPF ecosystem—making kernel observability tooling more accessible to Python developers.
Previous Talks:
- PyCon APAC 2024
- PyCon Hong Kong 2025
- PyCon Indonesia 2025
About this session
If you're a Python developer writing eBPF, you've probably used BCC. It embeds C code as strings in your Python scripts, then compiles them at runtime with Clang. Powerful, but slow to start.
In 2024, Python-BPF emerged. Instead of C strings, you write pure Python. The framework translates your AST directly into BPF bytecode via llvmlite.
I benchmarked both on the same tracepoint. BCC spent 93% of time in C compilation via Clang. Python-BPF spent 27% in Python AST translation. But Python-BPF's parser is rough—no ctx.args[0], no atomic operations.
I'll show the profiling data and help you decide which fits your use case.
For years, Python developers want to write eBPF programs had one dominant choice: BCC. BCC makes BPF programs easier to write, with kernel instrumentation in C (and includes a C wrapper around LLVM), and front-ends in Python and Lua. It is suited for many tasks, including performance analysis and network traffic control.
In 2024, Pragyansh Chaturvedi and Varun Mallya explored an alternative approach. They created Python-BPF, which translates pure Python AST into LLVM IR via llvmlite.
In this 30-minute session, I put both approaches under the microscope using time, strace, py-spy, and perf. I run identical tracepoint programs on both frameworks—capturing sys_enter_write, extracting PID and UID, and conditionally printing with bpf_printk—then compare what happens under the hood:
strace: BCC performs 2,342openat()calls (2,160 kernel headers) and 10,135read()calls during startup; Python-BPF forksllcas a child process with no kernel header traversalpy-spy: BCC spends 93% in C compilation via Clang versus Python-BPF's 27% in pure Python AST translationperf: Native call chain analysis revealing where BCC's Clang library calls block versus Python-BPF's direct IR pipelinetime -v: Startup latency comparison showing the cost of runtime compilation versus ahead-of-time generation
I then extend the comparison beyond tracepoints to other eBPF program types—revealing additional performance characteristics and parser limitations.
Finally, I confront Python-BPF's parser reality: ctx.args[0] breaks due to missing ast.Subscript support, and atomic map operations require atomicrmw instructions that don't yet exist.
Key Takeaways
- Why BCC's runtime compilation takes time (and when that trade-off is worth it)
- When Python-BPF's ahead-of-time approach fits your deployment
- The cognitive difference between writing C-in-strings and restricted Python syntax
Speakers
Alan is an Assistant Technical Manager at ATAL Engineering Group. He is specialised in the development of total solutions for smart buildings, covering the areas of energy optimisation, air conditioning, intelligent control, and automation. His recent work brings agentic AI into the facility management industry. Enabling AI reasoning in FM, and extending (semi-)automation into fuzzy, long-tail problems.
About this session
As we integrate LLMs into production chat applications, the challenge shifts from "can it answer?" to "can it answer reliably?" Pure text-generation is good for conversational engagement, but it often falls short when users need strict data formats, complex state management, or precise operational tasks.
This talk explores how to build robust, resilient agentic solutions by balancing trade-offs between LLM flexibility and system control. We will explore an operational robustness spectrum:
- Free Flow: Allowing the LLM to generate unstructured conversational responses, controlling systems using tools and sub-agents.
- Skills: Grounding the model with strict tool-calling capabilities.
- Adopting Steering: Utilizing modular prompting to guide tools and sub-agent behavior.
- Deterministic UI: The ultimate control mechanism—triggering embedded, generative UI components directly within the chat interface to guarantee predictability.
Beyond the agentic architecture, we will also tackle the testing dilemma. How do we validate a non-deterministic system? We’ll cover strategies for evaluating agentic workflows:
- from using pytest to verify tool execution
- , employing LLM-as-a-judge frameworks
- , to developing an internal testing platform for statistically managing pass rates.
Speakers
Adrian has been using Python for work for more than 20 years. Although not the only programming language to use, it is his favorite for quick experiments and prototyping. His interest is in mathematical modeling, number crunching, high performance computing, and in the last decade, machine learning and AI.
About this session
This talk shares the experience of building a Pythonic API out of a C++ library. The focus of this talk is to highlight the difference between a Python binding of a different language (usually C or C++ in reality), and a Pythonic library interface. Consequentially, you should not be satisfied with a binding.
The example I use in this talk is cuDNN, which you might already using it frequently but never realize that because it is hidden behind the other machine learning frameworks such as PyTorch. I will show you how you supposed to use cuDNN directly and therefore, why you probably prefer to use it via PyTorch instead. Then, I will show you my Pythonic API layer, which hides a lot of hassles to make your code cleaner and easier to maintain. Through this, I will explain, as a library creator, this is indeed more user-friendly and lowered the learning curve.
cuDNN is the cornerstone for frameworks like PyTorch but probably you never used it directly. cuDNN is a Linux library accessed via the C API. In recent versions, the preferred interface is a C++ header. While there is a Python binding to the C++ API, using it still feels like writing C or C++ code in Python.
To make cuDNN more accessible for Python developers, I've created a Pythonic interface that greatly reduces boilerplate. With this interface, you can focus on the core logic of your deep learning model, i.e., how to manipulate tensors, instead of managing tensor shapes, data types, or low-level CUDA kernel logistics.
Compared to the C++ API, this Pythonic interface highlights why Python is such an effective language to let human focus on the key elements of the code. In C++, you must set up multiple components before building a graph and handle manual validation and compilation before execution. Python's context manager mechanism provide a pathway to abstract these routines, making them seamless when creating a graph object. Python's duck-typing further increases flexibility in building computation graphs. Code examples will demonstrate that a good Python library is much more than just a direct binding to C++.
To illustrate the benefits for Python programmers, I'll demonstrate implementing a Llama model from scratch in PyTorch, but swapping out modules like nn.Linear, nn.RMSNorm, and nn.functional.scaled_dot_product_attention with their cuDNN-based drop-in replacements. While the results are functionally similar, you get fine-tuned implementations with significant performance gains. For example, using cuDNN for RMSNorm in Llama yielded a 10x speedup over PyTorch's implementation. By walking through the code, I'll show that using cuDNN directly is now straightforward for Python users.
While the example is about a deep learning model and PyTorch is used, the focus of this talk should be about how the API of other languages is different from Python, and you should not aim for a 1:1 mapping between the API of other language to Python (such as OpenCV) - since Python can do much better and you are obligated to make it more user-friendly.
Speakers
Dynamic and innovative Software Architect with over 10 years of experience delivering scalable, data-centric solutions across finance, telecom, and real estate industries. A passionate Python enthusiast and open-source contributor, I specialize in modern Python tooling, AI/LLM-powered applications, and building robust ETL pipelines and platform integrations. I thrive on leveraging cutting-edge technologies to solve complex business challenges, enhance operational efficiency, and create impactful solutions for both internal teams and external clients. With a strong focus on clean architecture, developer experience, and community collaboration, I actively contribute to open-source projects and enjoy sharing practical knowledge with the Python community.
About this session
You chose FastAPI because it promised speed. The benchmarks looked incredible. The docs were clean. Async was right there in the name. So why is your production service still crawling?
Here's the uncomfortable truth: FastAPI is fast — but only if you use it correctly. The framework gives you the tools; it doesn't protect you from yourself. After building and debugging FastAPI services across fintech, telecom, and real estate platforms — some handling hundreds of thousands of requests daily — I've seen the same patterns kill performance again and again. And most of them are invisible until your system is on fire.
This talk tears open the black box. We'll explore the five most dangerous anti-patterns that silently destroy FastAPI performance:
- Blocking the event loop: synchronous I/O hiding inside async route handlers — the single most common killer, and the hardest to catch in code review.
- Misconfigured dependency injection: Anthropic-style DI that looks clean but re-initializes heavyweight objects (DB connections, HTTP clients) on every request.
- N+1 queries through the ORM: FastAPI doesn't use Django's ORM, but SQLAlchemy lazy loading will bite you just as hard if you're not deliberate.
- Pydantic validation overhead: deeply nested models and
validatorchains that turn your serialization layer into your slowest layer. - Connection pool starvation: async DB drivers configured with defaults designed for development, quietly choking under real production concurrency.
For each anti-pattern, I'll show real profiling output, the exact code that caused it, and the fix — not a generic "use async properly" tip, but a concrete, copy-paste change you can apply to your codebase this week.
Attendees will leave with a personal FastAPI performance checklist, a profiling toolkit (py-spy, asyncio debug mode, SQLAlchemy echo), and a clear mental model of where FastAPI's async model actually buys you speed and where it doesn't.
If you're building production APIs with FastAPI — or planning to — this talk will save you weeks of debugging.
Background & Motivation
FastAPI has become the de facto Python web framework for API development, and for good reason: it offers native async support, automatic OpenAPI documentation, and Pydantic-backed validation. However, most articles and tutorials stop at "here's how to write an async route handler." The gap between a working FastAPI service and a fast FastAPI service is wide, filled with subtle traps that experienced engineers fall into routinely.
This talk is drawn from hands-on experience building and troubleshooting FastAPI services in high-throughput production environments — including payment gateways, data ingestion pipelines, and multi-tenant SaaS platforms.
Technical Depth
Anti-Pattern 1: Blocking the Event Loop
async def does not make your function non-blocking. If you call requests.get(),
time.sleep(), or any synchronous file I/O inside an async route, you freeze the
entire event loop. I'll demonstrate this with a live benchmark comparing:
- A route calling
requests.get()(sync) insideasync def - The same route using
httpx.AsyncClient - The same route using
run_in_executoras a migration bridge
Tools: asyncio debug mode, py-spy flame graphs.
Anti-Pattern 2: Dependency Injection Misuse
FastAPI's Depends() system is powerful but poorly understood. Developers often
initialize database sessions, HTTP clients, or config objects inside dependency
functions without understanding the lifecycle. I'll cover:
yield-based dependencies for proper resource managementlru_cache+@lru_cachewithfunctoolsfor singleton patterns- The
Annotated+Dependspattern (FastAPI 0.95+) for cleaner reuse
Anti-Pattern 3: SQLAlchemy Lazy Loading in Async Context
SQLAlchemy's async session (AsyncSession) does not support lazy loading — it
will raise MissingGreenlet errors or silently fall back to sync. This section
covers:
selectinloadvsjoinedloadin async contexts- When to use
lazy="raise"defensively in model definitions - Profiling with
SQLALCHEMY_WARN_20=1andecho=True
Anti-Pattern 4: Pydantic Overhead
Pydantic v2 is dramatically faster than v1, but misuse still costs you. I'll cover:
model_validatevs__init__performance difference- Avoiding redundant
.model_dump()/.model_validate()round-trips in middleware - Using
model_config = ConfigDict(from_attributes=True)correctly with ORM models - When to skip Pydantic entirely for internal models (dataclasses, TypedDict)
Anti-Pattern 5: Connection Pool Starvation
Most developers use create_async_engine with default settings: pool_size=5,
max_overflow=10. Under async concurrency, these defaults are dangerously low.
I'll show:
- How to benchmark pool exhaustion with
locust - Tuning
pool_size,max_overflow,pool_timeout, andpool_recycle - Using
asyncpgvsaiopgand why it matters
Demo Plan
All code examples will be drawn from a small but realistic "Orders API" — a multi-endpoint service backed by PostgreSQL, with realistic data volumes. The repo will be published on GitHub before the talk with profiling scripts included.
Profiling tools demonstrated:
py-spy(sampling profiler, zero-instrumentation)- Python's built-in
asynciodebug mode (PYTHONASYNCIODEBUG=1) cProfile+snakevizfor sync bottlenecks- SQLAlchemy's
echo+pool_logging_name
Talk Outline
[0:00 – 2:00] Opening hook: "Your async route isn't async" [2:00 – 5:00] Why FastAPI performance is misunderstood (brief async model recap) [5:00 – 10:00] Anti-pattern 1: Blocking the event loop (demo + fix) [10:00 – 14:00] Anti-pattern 2: Dependency injection lifecycle misuse (demo + fix) [14:00 – 18:00] Anti-pattern 3: SQLAlchemy lazy loading in async (demo + fix) [18:00 – 22:00] Anti-pattern 4: Pydantic validation overhead (demo + fix) [22:00 – 26:00] Anti-pattern 5: Connection pool starvation (demo + fix) [26:00 – 28:00] Performance checklist recap + profiling toolkit summary [28:00 – 30:00] Q&A
What Makes This Talk Different
Most "FastAPI tips" content is either surface-level (use async!) or too theoretical (here's how the event loop works). This talk sits in the practitioner's middle: real code, real profiling output, real fixes. Every pattern shown has been encountered in production — not constructed for the sake of a talk.
Speakers
Indy Ho is a registered physiotherapist and an Assistant Professor at the Technological and Higher Education Institute of Hong Kong. His research interests include sports science, sports therapy, strength and conditioning, data science, and machine learning. After obtaining his second master's degree in Data Science, he currently is studying in PhD to apply machine or deep learning for predictive analytics using IMU sensor data to realise the landing force and stability.
Ken Lee is a Technical Executive at Mosaic Digital Limited, a local AI and IT company.
Passionate about leveraging technology to benefit people, he specializes in AI Vision, AI Motion Analysis, Large Language Models (LLMs), and game development. With a strong focus on practical integration, Ken has built several innovative AI and motion-interactive solutions for NGOs and public exhibitions.
Beyond his executive leadership, he is an active technical writer dedicated to sharing his expertise in artificial intelligence and game technology with the broader developer community.
About this session
Python is currently widely used in solving research and application problems in the sports science, injury prevention, and physical fitness promotion areas. This talk will be divided into two sections. The first part will go through the existing literature regarding the current popular use of Python in solving different sports-related problems. Meanwhile, the first speaker will also show predictive analytics using machine/deep learning with Python scripts in biomechanical applications using IMU sensors and force plate data to help solve sports injury prevention problems. The second part will focus on the use of Python for motion capture and human movement analysis.
To collect human biomechanics data for sports and motion analysis, camera-based pose estimation with AI has become a popular alternative to Inertial Measurement Units (IMUs). While MediaPipe's pose tracking library may be the first one that comes up when you Vibe Code an application that needs human poses. But MediaPipe is just one option among many. Pose estimation libraries span from open-source to commercial, 2D and 3D output, prioritize either real-time inference speed or high-fidelity accuracy, and deployment complexity. In this section, we will explore the different approaches to pose estimation, how to choose the right one for your use case, and live demos of pose estimation in Python.
This part will include "The Landscape", as an exploration of different approaches to pose estimation, ranging from lightweight 2D skeletal tracking to more complex 3D kinematic modeling; "Choosing Your Stack" as a practical guide to evaluating trade-off; "hardware constraints, speed vs. accuracy, and ease of deployment" for showing the right library for your specific use case and; "Live Demos" for switching from slides to code, demonstrating pose tracking with YOLO-Pose. Here, participants can see how to track fast, complex physical mechanics with just a few lines of code.
This will be a great opportunity for us to exchange ideas and open the door to prepare for the future of sports, exercise, and human performance enhancement.
Speakers
Wenxin is currently a Ph.d. student in biostatistics at CityUHK. Her research interest lies in high-dimensional statistical machine learning and data science with applications in genetics and genomics data.
About this session
You know how to write tests. A statistician calls the same thing simulation. This talk is about what happens when those two worlds meet in Python.
As programmers, we trust code because tests pass. Statisticians trust code because it passes simulations: it recovers the truth on synthetic data, follows a specified distribution, and behaves well under stress scenarios. Same instinct, different vocabulary. In this talk I'll show what statisticians actually care about, why simulation is their test suite, and what that means for Python developers who want to make numerical code fast without leaving Python. We'll watch this play out on two workloads, an iterative one (k-means) and a massively repeated one (permutation tests), and see how choosing the right tool (NumPy, Numba, or JAX) comes down to the shape of the bottleneck.
Then I'll show it on a method from my own research: the same heavy numerical loop run across billions of independent problems. A handful of Numba kernels took it from "leave it running overnight" to "done over lunch", with NumPy-style code that collaborators could read and validate. No C++, no Rust.
There's a stereotype that Python is too slow for serious numerical work, so the moment things get heavy you must drop into C++ or rewrite in another language. Statisticians feel this acutely. Many rewrite their R hot paths in C++ via Rcpp and inherit deep low-level knowledge and brittle dependencies for the trouble. This talk argues there's a calmer path that stays in Python, and that the discipline for getting there is something every programmer already has, just under a different name.
Part 1 (about 20 min): From the tests you know to how statisticians think. I'll start where the audience already lives, automated testing, and map it onto statistical practice:
- A unit test becomes a fixed-seed equivalence check.
- A property test becomes a claim about behavior under the null, or about calibration.
- Stress and load tests become outliers, high dimension, and growing sample sizes.
- A regression test becomes "the fast version must still pass everything the reference version did."
That mapping unlocks the idea statisticians live by: speed follows validation. Real-world data has no answer key, so you simulate data where you already know the answer, write a readable NumPy reference, prove it recovers that answer, and only then optimize. The fast version is allowed in only if it answers the identical question. I'll make this concrete on two workloads with deliberately different bottleneck shapes:
- k-means, the iterative pattern: assign, update, repeat, where each step depends on the last. Tight loops, sequential state.
- Permutation test, the repeated-work pattern: shuffle, compute, repeat thousands of times. Independent, array-heavy.
For each, profiling reveals the shape of the bottleneck, and the shape picks the tool: BLAS-backed NumPy for dense algebra, Numba for numerical loops with no temporaries, and (a sentence each) threads for shared-array repetition and JAX or GPU for batched array programs. The point is the decision, not a tour of every tool.
Part 2 (about 10 min): Numba in a real research codebase. Then I'll use my own research code as a case study. The method is a heavy numerical loop run across billions of independent problems, and the original implementation was too slow to finish in a reasonable time. I rewrote the hot path in Numba, and it went from "leave it running overnight" to "done over lunch." You won't need any domain background. The only fact that matters is the compute shape.
- where the bottleneck actually was, not where I first guessed;
- keeping the readable NumPy version as the source of truth and adding
@njitkernels only on the proven hot path; - the refactors Numba demands, typed arrays, no fancy indexing, restructured loops, and what they bought in wall-clock time;
- why Numba beat a full rewrite here: tiny dependency footprint, code I could still debug, and code a collaborator could still read.
You leave with one transferable habit: write the readable version first, make simulation your test suite, profile to find the bottleneck's shape, and reach for the lightest tool that clears the bar, so complexity is spent only where it pays.
Speakers
劉凱晴小姐(Kazel Lau)畢業於香港大學(HKU),是香港領先的跨平台網絡安全與科技教育品牌HACKERTALE 的創辦人,同時擔任創新科技教育有限公司(ITED)的代表。作為一名經驗豐富的白帽駭客(Ethical Hacker)與資安技術專家,Kazel 專注於推動前沿 Agentic AI 與進攻性安全(Offensive Security)的技術融合,並積極在開源社群及各大科技論壇(如開源年會 HKOSCon、Open Data Day)分享前瞻視野與實戰工具。 在科技教育領域,Kazel 致力於填補 K12 階段的創科教學斷層。她巧妙地將網絡安全防護思維與機器人學相結合,透過 UBTECH UGOT AI 機械人生態系與開源 SDK 應用,引導中小學生從圖形化積木順暢銜接至真實的 Python 文字編程。Kazel 期盼透過將硬體控制、邊緣運算(Edge AI)與硬核技術普及化,為香港培育具備真實除錯思維與未來國際競爭力的創科新血。
About this session
在 K-12 程式教育中,學生從圖形化積木(Blockly)過渡至純文字 Python 時常面臨嚴峻的教學斷層,學生往往因枯燥的語法除錯(Syntax debugging)而失去對程式邏輯的興趣。本場演講提出一種創新的教學路徑,探討如何藉由開源的 ugot Python SDK 與 UBTECH UGOT 硬體生態系,透過「硬體反饋」與「雙模並行代碼提示」突破此教學瓶頸。我們將深入解析 UGOT 主控晶片內建 NPU 的技術架構,展示其如何在本地端(Edge AI)高速執行電腦視覺演算法(如 Teachable Machine、Apriltag 標籤追蹤、人臉偵測)。同時,本研究聚焦於 uCode 平台如何實時將 Blockly 積木編譯為標準 Python 3 代碼,使學生在調整麥克納姆輪全向移動或多關節四足步態等運動參數時,能直觀理解函數呼叫、物件導向(OOP)與異步數據處理。最後,本演講將結合全球青少年機器人競賽(如 Robo Genius)實例,分析學生如何利用 Python 的模組化特性,在多形態機器人切換時快速重構代碼,進而在競賽壓力下建立紮實的實戰除錯思維。
詳細介紹 (Description) :
在 K12 程式教育中,從圖形化積木(Blockly)過渡到純文字 Python 一直存在巨大的教學斷層。學生往往卡在枯燥的語法除錯(Syntax debugging),而失去了對程式邏輯的興趣。本場演講將分享如何利用開源的 ugot Python SDK 與 UBTECH UGOT 硬體生態系,透過「硬體反饋」與「雙模並行代碼提示」來打破這個瓶頸。
我們將深入探討 UGOT 的技術架構,包括其主控晶片如何透過內置 NPU 在本地端(Edge AI)執行高速的電腦視覺演算法(如 Apriltag 標籤追蹤、Teachable Machine 、人臉偵測、顏色與線條識別)。演講重點將放在 uCode 平台如何實時將 Blockly 積木編譯為標準的 Python 3 代碼。透過這種「改動積木,即時對照 Python 語法」的動態反饋,學生能在調整機械人運動參數(如麥克納姆輪全向移動、多關節四足步態)的同時,直觀地理解 Python 的函數呼叫、物件導向(OOP)概念以及實時傳感器數據的異步處理。
最後,我們將結合近年本地及全球青少年機械人競賽(如 Robo Genius)的實際案例,分析學生在面對多形態機械人(如平衡車、工程車、蜘蛛形態)切換時,如何利用 Python 的模組化特性快速重構代碼,並在競賽壓力下建立真實的 Python 除錯思維。
核心收穫 (Key Takeaways):
硬體驅動的語法過渡: 掌握如何利用 uCode 的雙模代碼引擎,透過即時 Python 語法提示與硬件即時反饋,降低 K12 學生進入文字編程的焦慮感。
Edge AI 與 ugot SDK 實踐: 了解 UGOT 核心 SDK 的底層邏輯,看 Python 如何在不依賴雲端的情況下,調用本地 NPU 進行多種視覺與語音辨識任務。
競賽推動的實戰思維: 獲取源自真實機械人競賽的課程設計與除錯教學策略,看學生如何利用 Python 應對多形態(Multi-mimetic)硬件結構的控制挑戰。
Speakers
- PyLadies Seoul Organizer
- Django Korea Organizer
- AWS Women in Cloud Organizer
- Software engineer at SocraAI (formerly Riiid), an AI edtech company
- Pythonista using Python and Django
- Experienced speaker at PyCon
- Proud owner of a cute dog
- Korean, mainly communicating in English, and learning a little Japanese.
About this session
Language barriers can be a challenge for non-native developers aiming for the global stage. In PyLadies Seoul, we tackled this problem using Python. I built an automated AI tutor that transcribes spoken English using an STT API and generates detailed feedback via an LLM.
This talk is not just about an English study group. It’s a technical walkthrough of how I built this AI backend using Python. Processing audio and waiting for LLM responses involves handling long-running tasks, managing states, and dealing with API failures. I will share the architecture behind this AI pipeline, the engineering challenges of managing heavy API workflows, and how you can apply this architecture to solve other everyday problems like meeting summaries or automated reports.
Speakers
Clarissa Gunawan is a data workflow enthusiast who believes that messy processes lead to missed opportunities. At Bloomberg, she is part of the Fixed Income Data team in Hong Kong, where she transforms workflows into clean, efficient systems using Python. With a master's degree in business analytics from Carnegie Mellon University, Clarissa brings structure to chaos and ensures that data can drive real impact, when and where it is needed.
About this session
We've all been through the process of writing and rewriting Python code for personal projects until it works as expected, only to realise that it has become an unreadable mess. The next thing you know, you no longer understand the code you've written.
In this talk, I will share my personal journey of feeling overwhelmed by Python's styling guidelines. That is, until I discovered Ruff, a lightning-fast open source tool that did not just fix my code, but actually taught me how to write it better.
We will discuss how Ruff can serve as a Python mentor for learners by:
- Improving code clarity through its instant feedback on how and why code can be improved
- Improving standardisation to align across projects and drastically speed up the code refactoring process
- Enforcing highly customisable rules depending on your project needs
Whether you are a beginner looking to adopt and follow Python's best practices, or an expert looking to standardise your workflow across projects, you will learn how Ruff can help as your personal coding companion.



































