COSCUP x UbuCon Asia 2026 Logo
  • About
  • Session
  • Transportation
  • Venue
  • Communities
  • Sponsors
    • Participate Guide
    • First Timer
    • Activity
    • Speaker Participation
    • Welcome Party
    • For Overseas Visitors
    • Open Source Communities
    • Sponsorship Partners
    • Invitation Letter Guide
  • Staff
  • Fringe Events / BoF
  • Blog
  • CoC
中文

Hybrid AI: The Next Pattern for AI-Powered Apps

Time
2026-08-09 11:45 ~ 12:15
Speaker
Sasha Denisov
Room
AU
Co-write
https://hackmd.io/r1udJ4e8fe
Feedback
https://coscup.org/2026-feedback/?session=NJSSSE
Open LLM End User: Downstream Applications IntermediateEnglish

Abstract

Most applications treat AI as a cloud-only feature — send a request, wait for a response, pay per token. Open models like Gemma, Llama, DeepSeek, and Phi freed developers from proprietary APIs, but not from costs: you still need to host them somewhere — GPU servers, infrastructure, per-request pricing. You can run models directly on the user's device — zero inference costs, offline access, full data privacy. But there's a trade-off: models small enough to run on a phone or in a browser can't handle every task. Hybrid AI is the answer to this dilemma. Combine cloud and on-device models in a single application: simple and privacy-sensitive tasks run on-device, complex reasoning and multimodal workloads go to the cloud. The problem is that you end up maintaining two completely different inference stacks: different runtimes (TFLite, LiteRT-LM, llama.cpp), different model formats (TFLite flatbuffers vs GGUF), different quantization strategies, and different streaming APIs — all within one app. In this talk, I'll show how to solve this problem with Genkit — an open-source AI framework from Google with an open plugin system. Each plugin adapts a specific runtime or model format to a unified pipeline: genkit_flutter_gemma for TFLite/LiteRT models, genkit_llamadart for GGUF. Engine-specific details are hidden behind a single API for flows, structured output, tool calling, and agentic workflows. Need a new runtime? Write a plugin. New model format? Same story. On-device inference runs across Android, iOS, macOS, Windows, Linux, and Web — switching between cloud and local execution is a one-line model reference change. We'll explore hybrid patterns in practice: how to route between on-device and cloud based on model capabilities, connectivity, and task complexity. We'll cover the real trade-offs — quantization impact on output quality, memory constraints across platforms, cold start latency, and when hybrid actually makes sense versus pure cloud or pure edge. Open-source models made on-device AI possible. Open-source tooling makes hybrid AI extensible.

Speaker

Sasha Denisov

Sasha Denisov

Sasha is CTO at Brainform.ai with over 20 years of experience architecting scalable enterprise systems. With a strong engineering background, his expertise spans frontend, backend, cloud infrastructure, mobile development, and AI — from cloud-based generative AI to on-device solutions. He specializes in building robust, production-ready products using a variety of technologies and frameworks. Sasha has delivered solutions across fintech, digital media, and entertainment. He is a Google Developer Expert for Cloud, AI, Firebase, Flutter, and Dart, co-organizes the Flutter Berlin Community, and is a recognized international speaker and writer, having presented at 30+ conferences worldwide.

Diamond

Cumulative4 YrsCanonical

Gold

Consecutive2 YrsInformation Management AssociationConsecutive2 YrsCathay Financial HoldingsConsecutive5 YrsE.SUN BankCumulative13 YrsMySQLCumulative6 YrsBerry AINitra

Silver

Cumulative16 YrsGamesofa Inc.Cumulative3 YrsLinux Professional Institute

Bronze

Collaborator2 YrsKKTIXCumulative3 YrsNational Center for High-performance ComputingONLYOFFICECumulative6 YrsQNAP Systems, Inc.Consecutive16 YrsThe Archilife Research FoundationSUN SQUARE Co., LtdPenpeer

Friend

Consecutive3 YrsAndes Technology CorporationCumulative12 YrsAppierCumulative3 YrsGeode Labs & ETHTaipei

Special Thanks

Consecutive2 YrsDepartment of Information Technology, Taipei City GovernmentCollaborator3 YrsSiFive數位發展部Rozeta AIGrafana LabsLive YoungCollaborator5 YrsThree Wheel Nitro CoffeeCollaborator9 YrsHackmdCollaborator4 Yrs天瓏書局Collaborator3 YrsTapPayIndividual-Sponsorship

Co-host

Collaborator9 YrsNTUST - Department of Electronic and Computer Engineering

Co-organizer

Collaborator12 YrsOpen Culture Foundation

COSCUP x UbuCon Asia 2026

Conference for Open Source Coders, Users, and Promoters | Asia's largest open source community conference.

Contact

  • Attendee Services
  • Sponsorship
  • Call for Proposals
  • Marketing

Resources

  • COSCUP Blog
  • Newsletter Subscription
  • Event Photos

Sitemap

  • Home
  • About
  • Transportation
20062007200820092010201120122013201420152016201720182019202020212022202320242025