COSCUP x UbuCon Asia 2026 標誌
  • 關於我們
  • 議程
  • 交通
  • 會場地圖
  • 社群
  • 贊助夥伴
    • 參與指南
    • 第一次參與
    • 活動參與
    • 講者參與
    • 前夜派對
    • 海外參與者
    • 開源社群、攤位及議程軌
    • 贊助夥伴
    • 邀請函申請指南
  • 工作人員
  • 周邊活動 / BoF
  • 部落格
  • 社群守則
English

Hybrid AI: The Next Pattern for AI-Powered Apps

時間
2026-08-09 11:45 ~ 12:15
講者
Sasha Denisov
位置
AU
共筆
Open LLM End User: 開源模型應用 中階英語

議程簡介

Most applications treat AI as a cloud-only feature — send a request, wait for a response, pay per token. Open models like Gemma, Llama, DeepSeek, and Phi freed developers from proprietary APIs, but not from costs: you still need to host them somewhere — GPU servers, infrastructure, per-request pricing. You can run models directly on the user's device — zero inference costs, offline access, full data privacy. But there's a trade-off: models small enough to run on a phone or in a browser can't handle every task. Hybrid AI is the answer to this dilemma. Combine cloud and on-device models in a single application: simple and privacy-sensitive tasks run on-device, complex reasoning and multimodal workloads go to the cloud. The problem is that you end up maintaining two completely different inference stacks: different runtimes (TFLite, LiteRT-LM, llama.cpp), different model formats (TFLite flatbuffers vs GGUF), different quantization strategies, and different streaming APIs — all within one app. In this talk, I'll show how to solve this problem with Genkit — an open-source AI framework from Google with an open plugin system. Each plugin adapts a specific runtime or model format to a unified pipeline: genkit_flutter_gemma for TFLite/LiteRT models, genkit_llamadart for GGUF. Engine-specific details are hidden behind a single API for flows, structured output, tool calling, and agentic workflows. Need a new runtime? Write a plugin. New model format? Same story. On-device inference runs across Android, iOS, macOS, Windows, Linux, and Web — switching between cloud and local execution is a one-line model reference change. We'll explore hybrid patterns in practice: how to route between on-device and cloud based on model capabilities, connectivity, and task complexity. We'll cover the real trade-offs — quantization impact on output quality, memory constraints across platforms, cold start latency, and when hybrid actually makes sense versus pure cloud or pure edge. Open-source models made on-device AI possible. Open-source tooling makes hybrid AI extensible.

講者

Sasha Denisov

Sasha Denisov

Sasha is CTO at Brainform.ai with over 20 years of experience architecting scalable enterprise systems. With a strong engineering background, his expertise spans frontend, backend, cloud infrastructure, mobile development, and AI — from cloud-based generative AI to on-device solutions. He specializes in building robust, production-ready products using a variety of technologies and frameworks. Sasha has delivered solutions across fintech, digital media, and entertainment. He is a Google Developer Expert for Cloud, AI, Firebase, Flutter, and Dart, co-organizes the Flutter Berlin Community, and is a recognized international speaker and writer, having presented at 30+ conferences worldwide.

鑽石級

Canonical

黃金級

連續贊助2 年中華民國資訊經理人協會連續贊助2 年國泰金融控股公司連續贊助5 年玉山銀行累計贊助12 年MySQL累計贊助6 年華捷智能股份有限公司Nitra

白銀級

累計贊助16 年慧邦科技股份有限公司  Gamesofa Inc.

青銅級

華娛網路娛樂股份有限公司財團法人國家實驗研究院國家高速網路與計算中心ONLYOFFICEQNAP Systems, Inc. 威聯通科技連續贊助16 年祐生研究基金會源鋼技術顧問有限公司 SUN SQUARE Co., Ltd

好朋友級

連續贊助3 年晶心科技股份有限公司沛星互動科技股份有限公司

特別感謝

連續贊助2 年臺北市政府資訊局美商賽發馥股份有限公司臺灣分公司Rozeta AI

共同主辦單位

累計合作9 年臺灣科技大學 電子工程系

協辦單位

累計合作12 年開放文化基金會

COSCUP x UbuCon Asia 2026

Conference for Open Source Coders, Users, and Promoters | 亞洲最大開源年會,由社群自發舉辦。

聯絡我們

  • 會眾服務
  • 贊助合作
  • 議程投稿
  • 行銷方案

相關資源

  • COSCUP 部落格
  • 訂閱電子報
  • 活動照片

網站導覽

  • 首頁
  • 關於我們
  • 交通資訊
  • 贊助夥伴
20062007200820092010201120122013201420152016201720182019202020212022202320242025