Skip to main content Start main content

PolyU GenAI App updates on 1st September 2026

GPT-5.6

 

OpenAI released the GPT-5.6 family of models on 9 July 2026, including its new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, its most cost-efficient model. GPT-5.6 supports up to 1 million tokens of context, with training data covers information up to June 2026.

GPT‑5.6 sets a new standard for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science, while outperforming previous and competing frontier models with fewer tokens and at a lower estimated cost. The result is stronger performance per dollar : more successful work for the same spend, or comparable results at a lower total cost. 

GPT-5.6 Sol delivers the most advanced reasoning capabilities yet, supporting extended reasoning, agentic workflows, and code-focused scenarios, for the most demanding enterprise workloads.

The PolyU Gen AI app released GPT-5.6 Sol on 1 September 2026, replacing GPT-5.5 (2026-04-24).

GPT-5.6 Sol, hosted on Azure Cloud and accessed via the PolyU Gen AI app, will consume credits as follows:

  • <= 272K input tokens: 5 credits per 1K tokens for text or image input, and 30 credits per 1K tokens for output text, including chain-of-thought tokens.
  • > 272K input tokens: 10 credits per 1K tokens for text or image input, and 45 credits per 1K tokens for output text, including chain-of-thought tokens.

 

GPT-5.6 Terra is a balanced model for everyday work, delivering performance competitive with GPT-5.5 at a lower cost, making it ideal for scaling intelligent applications across the enterprise.

The PolyU Gen AI app released GPT-5.6 Terra on 1 September 2026, replacing GPT-5.4 (2026-03-05).

GPT-5.6 Terra, hosted on Azure Cloud and accessed via the PolyU Gen AI app, will consume credits as follows:

  • <= 272K input tokens: 2 credits per 1K tokens for text or image input, and 12 credits per 1K tokens for output text, including chain-of-thought tokens.
  • > 272K input tokens: 4 credits per 1K tokens for text or image input, and 18 credits per 1K tokens for output text, including chain-of-thought tokens.

 

GPT-5.6 Luna is the fastest and most affordable model in the family, making it well suited to high-volume, latency-sensitive workloads.

The PolyU Gen AI app released GPT-5.6 Luna on 1 September 2026, replacing GPT-5.4-mini (2026-03-17).

GPT-5.6 Luna, hosted on Azure Cloud and accessed via the PolyU Gen AI app, will consume credits as follows:

  • <= 272K input tokens: 0.2 credits per 1K tokens for text or image input, and 1.2 credits per 1K tokens for output text, including chain-of-thought tokens.
  • > 272K input tokens: 0.4 credits per 1K tokens for text or image input, and 1.8 credits per 1K tokens for output text, including chain-of-thought tokens.

 

For more details on GPT-5.6, please refer to Introducing GPT5.6 and OpenAI's GPT-5.6 in Microsoft Foundry.

 

 

 

Gemini-3.7-Flash

02

Google Cloud released Gemini 3.7 Flash, its most intelligent workhorse model yet for coding and agents, on 13 August 2026.  This release came just three weeks after Gemini 3.6 Flash and was a direct result of developer feedback and algorithmic innovations that Google looks forward to bringing to future models.

Gemini 3.7 Flash delivers substantial improvements across software engineering, knowledge work, and web development workflows — with an introductory price at half the original 3.6 Flash cost per million tokens.

The PolyU Gen AI app released the Gemini 3.7 Flash model on 1 September 2026, replacing Gemini 3.6 Flash.

Gemini 3.7 Flash, hosted on Google Cloud and accessed via the PolyU Gen AI app, consumes 0.75 credits per 1K tokens for text or image input and 3.75 credits per 1K tokens for text output, including chain-of-thought tokens.

For details of Gemini 3.7 Flash, please refer to the Gemini 3.7 Flash Official Blog and the Gemini 3.7 Flash Model Card.

 

 

Kimi-K3

03

Moonshot AI released Kimi K3 on 16 July 2026. Kimi K3 is Kimi’s most capable flagship model to date, with 2.8 trillion parameters. It is built on Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals, with native visual understanding and a 1M-token context window. It is the world’s first open-source model in the 3-trillion-parameter class, designed for frontier intelligence scenarios including long-horizon coding, knowledge work, and reasoning.

While its overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol, Kimi K3 demonstrated frontier-level performance across Moonshot’s evaluation suite, consistently outperforming other tested models.

The PolyU Gen AI app introduced Kimi K3 on 1 September 2026, replacing Kimi K2.6.

The Kimi K3 model, hosted on Alibaba Cloud Bailian and accessible via the PolyU Gen AI app, consumes 3 credits per 1K tokens for text or image input and 15 credits per 1K tokens for text output, including chain-of-thought reasoning.

For more details on Kimi K3, please refer to the Kimi K3 Technical Blog and the Research Paper.

 

 

Qwen3.8-Max

04

Alibaba Cloud released Qwen3.8-Max, the most capable model in the Qwen family to date,  on 3 August 2026. Built upon the architectural foundation of Qwen 3.5, Qwen 3.8-Max scales to 2.4 trillion parameters, delivering comprehensive improvements across coding, work, research, and long-horizon tasks. It can not only answer more challenging questions, but also complete complex tasks end-to-end with greater reliability, producing dependable deliverables.

The PolyU Gen AI app released Qwen3.8-Max on 1 September 2026, replacing Qwen3.7-Plus (2026-05-26).

Qwen3.8-Max, hosted on Alibaba Cloud Bailian and accessed via the PolyU Gen AI app, consumes 2 credits per 1K tokens for text or image input, and 6 credits per 1K tokens for text output, including chain-of-thought tokens.

For details on Qwen3.8-Max, please refer to the Qwen3.8-Max Official Blog.

 

 

DeepSeek-V4

DeepSeek

DeepSeek AI introduced DeepSeek-V4-Pro-0813 on 13 August 2026 and DeepSeek-V4-Flash-0731 on 31 July 2026. These models are the official release of DeepSeek-V4, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments.

The PolyU Gen AI app introduced DeepSeek-V4-Pro-0813 on 1 September 2026, replacing DeepSeek-V4-Pro.

The DeepSeek-V4-Pro-0813 service, powered by Alibaba Cloud Bailian, consumes 1.272 credits per 1,000 tokens for input and 3.816 credits per 1,000 tokens for output, plus the cost of chain-of-thought tokens.

The PolyU Gen AI app introduced DeepSeek-V4-Flash-0731 on 1 September 2026, replacing DeepSeek-V4-Flash.

The DeepSeek-V4-Flash-0731 service, powered by Alibaba Cloud Bailian, consumes 0.424 credits per 1,000 tokens for input and 1.272 credits per 1,000 tokens for output, plus the cost of chain-of-thought tokens.

For more information, please refer to the DeepSeek-V4-Pro-0813 Official Repository and DeepSeek-V4-Flash-0731 Official Repository.

 

 

Your browser is not the latest version. If you continue to browse our website, Some pages may not function properly.

You are recommended to upgrade to a newer version or switch to a different browser. A list of the web browsers that we support can be found here