
Gemini 3.6 Flash
Google Cloud released Gemini 3.6 Flash on 21 July 2026. Gemini 3.6 Flash builds directly on developer and customer feedback from 3.5 Flash. 3.6 Flash not only delivers a step up in coding and knowledge work, but also improves token efficiency. It also requires fewer reasoning steps and tool calls to accomplish multi-step workflows. The model’s training data covers information up to March 2026.
The PolyU Gen AI app released the Gemini 3.6 Flash model on 1 August 2026, replacing Gemini 3.5 Flash.
Gemini 3.6 Flash, hosted on Google Cloud and accessed via the PolyU Gen AI app, consumes 1.5 credits per 1K tokens for text or image input and 7.5 credits per 1K tokens for text output, including chain-of-thought tokens.
For details on Gemini 3.6 Flash, please refer to the Introduction to Gemini 3.6 Flash, the Gemini 3.6 Flash Official Blog and the Gemini 3.6 Flash Model Card.
Gemini 3.5 Flash-Lite
Google Cloud released Gemini 3.5 Flash-Lite on 21 July 2026. Gemini 3.5 Flash-Lite is designed for both low-latency tasks and tasks where high throughput is critical, such as agentic search and document processing. The model’s training data covers information up to March 2026.
Gemini 3.5 Flash-Lite outperforms 3.1 Flash-Lite. Depending on the workload, developers can configure the model to prioritize low-latency, low-cost execution for high-volume tasks with minimal and low thinking levels, or engage higher thinking levels to process multi-step sub-agent workloads.
The PolyU Gen AI app released the Gemini 3.5 Flash-Lite model on 1 August 2026.
Gemini 3.5 Flash-Lite, hosted on Google Cloud and accessed via the PolyU Gen AI app, consumes 0.3 credits per 1K tokens for text or image input and 2.5 credits per 1K tokens for text output, including chain-of-thought tokens.
For details on Gemini 3.5 Flash-Lite, please refer to the Introduction to Gemini 3.5 Flash-Lite, Gemini 3.5 Flash-Lite Official Blog and the Gemini 3.5 Flash-Lite Model Card.

GLM-5.2
Z.ai introduced GLM-5.2 on 16 June 2026. GLM-5.2's new capabilities include:
- Solid 1M Context: A solid 1M-token context window that sustains long-context work.
- Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking-effort levels to balance performance and latency.
- Improved Architecture: Z.ai reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9× at a 1M context length. It also improved GLM-5.2’s MTP layer for speculative decoding, increasing the acceptance length by up to 20%.
The GLM-5.2 model was released on the PolyU Gen AI app on 1 August 2026 for use by PolyU staff and students.
The GLM-5.2 service, powered by Alibaba Cloud Bailian, incurs a usage fee of 1.12 credits per 1,000 tokens for input and 3.92 credits per 1,000 tokens for output, plus the cost of chain-of-thought tokens.
For more information, please refer to the GLM-5.2 Official Repository, the GLM-5.2 Official Blog and Prompt Engineering for OpenAI’s Reasoning Models.
MiniMax M3
MiniMax AI released MiniMax M3 on 1 June 2026. M3 uses MSA (MiniMax Sparse Attention), a new attention architecture proposed by its team, and supports ultra-long context windows of up to 1M tokens. MiniMax M3 is a natively multimodal model that supports image and video input.
The PolyU Gen AI app introduced MiniMax M3 on 1 August 2026.
The MiniMax M3 model, hosted on Alibaba Cloud Bailian and accessible via the PolyU Gen AI app, consumes 0.084 credits per 1K tokens for text or image input and 0.24 credits per 1K tokens for text output, including chain-of-thought tokens.
For more details on MiniMax-M3, please refer to the MiniMax M3 Official Blog and the Official Hugging Face Repository.
Kimi K2.6
Moonshot AI released Kimi K2.6 on 20 April 2026. Kimi K2.6 is a multimodal agentic model with the following key features:
- Long-Horizon Coding: K2.6 achieves significant improvements in complex, end-to-end coding tasks.
- Coding-Driven Design: K2.6 can transform simple prompts and visual inputs into production-ready interfaces and lightweight full-stack workflows, generating structured layouts, interactive elements, and rich animations.
- Elevated Agent Swarm: K2.6 can dynamically decompose tasks into parallel, domain-specialized subtasks, delivering end-to-end outputs ranging from documents and websites to spreadsheets in a single autonomous run.
- Proactive & Open Orchestration: For autonomous tasks, K2.6 demonstrates strong performance in powering 24/7 background agents that proactively manage schedules, execute code, and orchestrate cross-platform operations without human oversight.
The PolyU Gen AI app introduced Kimi K2.6 on 1 August 2026, replacing Kimi K2.5.
The Kimi K2.6 model, hosted on Alibaba Cloud Bailian and accessed via the PolyU Gen AI app, consumes 0.91 credits per 1K tokens for text or image input and 3.78 credits per 1K tokens for text output, including chain-of-thought reasoning.
For more details on Kimi K2.6, please refer to the Kimi K2.6 Technical Blog and the Official Hugging Face Repository.
Tencent Hy3
Tencent released and open-sourced the Hy3 model on 6 July 2026. Following the Hy3 preview launch in late April, Tencent gathered feedback from preview model users and scaled up post-training with higher-quality data.
The Tencent Hy3 model was released on the PolyU Gen AI on 1 August 2026, replacing the Tencent Hy3 preview.
Tencent Hy3, powered by Tencent Cloud, and accessed via the PolyU Gen AI app, incurs a usage fee of 0.14 credits per 1K tokens for input text and 0.56 credits per 1K tokens for output text, including chain-of-thought tokens.
For more information, please refer to the Tencent Hy3 Official Blog, the Official Hugging Face Repository and Prompt Engineering for OpenAI’s Reasoning Models.