Skip to main content
AI & Models10 min read

Million-Token Contexts Go Mainstream: GLM-5.2, Fugu Ultra, Nemotron 3 Ultra, and New 262K Coding Models

This week’s model releases show the long-context race accelerating, with multiple 1M-token models arriving alongside open-weight coding models built for repository-scale understanding. GLM-5.2, Fugu Ultra, NVIDIA Nemotron 3 Ultra, Laguna XS 2.1, and Kimi K2.7 Code point toward AI systems that can reason over entire codebases, research corpora, and extended technical workflows—but context length alone still does not guarantee reliability.

The biggest theme in this week’s AI model releases is not just bigger models—it is bigger working memory. Multiple new or recently released models now advertise context windows around one million tokens, while the coding-focused entrants are settling into the 262K-token range, enough to handle substantial repositories, documentation sets, and multi-file engineering tasks in a single prompt.

That shift matters because long-context models change the shape of practical AI use: instead of carefully excerpting a few files or documents, users can increasingly ask models to reason across whole systems. The caveat is equally important: large context windows improve access to information, but they do not automatically solve retrieval accuracy, instruction following, latency, cost, or hallucination.

ModelProviderContextPricingKey Capabilities
GLM-5.2Zhipu AI / Z.ai1,048,576 tokensNot specified in release data; check OpenRouter/Ollama provider listingLong-context reasoning, text generation, coding
Fugu UltraSakana AI1,000,000 tokensNot specified in release data; check OpenRouter listingLong-context reasoning, general-purpose text generation
NVIDIA Nemotron 3 Ultra 550B A55BNVIDIA1,000,000 tokensNot specified in release data; check OpenRouter/Ollama provider listingLong-context reasoning, text generation, coding
Laguna XS 2.1Poolside262,144 tokensNot specified in release data; check OpenRouter/Ollama provider listingCode generation, long-context text workflows
Kimi K2.7 CodeMoonshot AI262,144 tokensNot specified in release data; check OpenRouter/Ollama provider listingSoftware engineering, code understanding, reasoning

GLM-5.2: an open-weight 1M-token foundation model for reasoning, coding, and generation

GLM-5.2 from Zhipu AI / Z.ai is one of the most notable releases in this batch because it combines a 1,048,576-token context window with open-weight availability. That puts it in a category that is still relatively rare: very-long-context models that can be run or integrated outside a purely closed API environment, depending on the deployment route and hardware constraints.

The model targets long-context reasoning, general text generation, and code-generation workloads. In practical terms, that means it is positioned for tasks such as reading large technical manuals, comparing many documents, analyzing large codebases, or maintaining continuity across extended agentic workflows. The extra context headroom is especially relevant for tasks where the answer depends on scattered details rather than a single passage.

Technically, the headline specification is the 1,048,576-token context length. The model is listed on OpenRouter and is also available through the Ollama library. It is open weight, though users should still verify the exact license terms, acceptable-use constraints, quantization options, and hardware requirements from the provider or distribution channel. The release data does not specify maximum output length, multimodal support, or pricing, so those should be treated as deployment-specific details rather than assumed capabilities.

The main benefit of GLM-5.2 is flexibility. Open-weight access makes it attractive for teams that want more control over inference environments, privacy boundaries, or customization pipelines. The long context also makes it a strong candidate for retrieval-light workflows where users want to provide the full source material directly.

The trade-off is that million-token prompting can be expensive and slow, even when a model supports it. Long-context reliability is also uneven across the industry: models may attend well to recent or prominent sections while missing details buried in the middle. GLM-5.2 should therefore be evaluated not just on maximum context length, but on needle-in-a-haystack retrieval, citation fidelity, coding correctness, and latency under realistic workloads.

Compared with smaller or shorter-context open models, GLM-5.2’s differentiator is scale of context. Compared with closed long-context systems, its appeal is likely to be deployment control and openness rather than guaranteed turnkey performance.

Fugu Ultra: Sakana AI enters the 1M-token general-reasoning tier

Fugu Ultra is Sakana AI’s long-context model listed on OpenRouter with a 1,000,000-token context window. Unlike the coding-specific releases in this roundup, Fugu Ultra is positioned as a general-purpose reasoning and text-generation model, making it relevant for broad analytical workflows rather than only software engineering.

The notable feature is straightforward: one million tokens of context. That is enough to ingest book-length material, extensive legal or technical corpora, long chat histories, or multi-document research packets. For users building research assistants, document analysts, or complex planning agents, that context size can reduce the need to aggressively chunk and summarize every input before the model sees it.

Its listed capabilities are long-context reasoning and text generation. The release data does not indicate open-weight availability; Fugu Ultra is marked as not open weight. It is available through OpenRouter, but maximum output length, modalities, and pricing are not specified in the provided release information. Readers should verify current rates and operational limits in the live provider listing before budgeting large-context workloads.

Fugu Ultra’s strength is its positioning as a broad reasoning model rather than a specialized coding assistant. If Sakana’s implementation handles long-range attention robustly, it could be useful for synthesis tasks where the model must compare arguments, track timelines, or reconcile details across many documents.

The obvious limitation is access and control. Because it is not open weight, users depend on hosted availability, provider terms, and API economics. Another caveat is that long context can create a false sense of completeness: just because a million tokens are present does not mean the model will use all of them equally well. For high-stakes tasks, users should still require grounded citations, intermediate reasoning checks, and targeted retrieval.

Relative to GLM-5.2 and Nemotron 3 Ultra, Fugu Ultra’s biggest distinction is that it is a closed, general-purpose long-context option. It may appeal to users who value managed access over local deployment, but it will need strong empirical performance to stand out in an increasingly crowded 1M-token field.

NVIDIA Nemotron 3 Ultra 550B A55B: a heavyweight open model with million-token reach

NVIDIA Nemotron 3 Ultra 550B A55B is another major entry because it brings the Nemotron family into the million-token context tier while remaining open weight. The model is listed on OpenRouter and available in the Ollama library, giving it visibility across both hosted routing and local-model ecosystems.

The name indicates a very large 550B-class model, with the “A55B” designation suggesting an active-parameter profile typical of sparse or mixture-style architectures, though users should confirm the exact architecture in NVIDIA’s technical materials. Its intended workloads include long-context reasoning, general text generation, and code generation.

The core specification is a 1,000,000-token context window. Like the other models in this roundup, the provided release data does not specify maximum output length, pricing, or non-text modalities. It is open weight, but license details and practical deployment requirements matter substantially here: a model of this class may be open, but not necessarily easy or inexpensive to run at full capability.

The biggest strength of Nemotron 3 Ultra is that it combines scale, openness, and a broad capability profile. NVIDIA’s models are often evaluated by enterprise and infrastructure-focused users who care about throughput, deployment control, and integration with accelerated inference stacks. For organizations already invested in NVIDIA hardware, Nemotron-family models may be particularly interesting.

The limitations are practical. Large open-weight models can impose significant hardware, memory, quantization, and serving complexity. Million-token context also magnifies those costs. Users should benchmark smaller prompts, long-context retrieval, coding accuracy, and concurrency before assuming that the maximum context window is economically usable in production.

Compared with GLM-5.2, Nemotron 3 Ultra appears positioned more as a heavyweight infrastructure model. Compared with Fugu Ultra, its open-weight status is the major differentiator.

Laguna XS 2.1: Poolside’s 262K-token model for long-context coding and text

Laguna XS 2.1 from Poolside is smaller in context than the million-token models above, but its 262,144-token window is still substantial—especially for software work. It is listed on OpenRouter and available in the Ollama library, and it is open weight.

The model is aimed at code generation, long-context workflows, and text generation. A 262K-token context can often fit a meaningful slice of a repository, including source files, tests, configuration, documentation, and recent issue context. For many coding tasks, that may be more practical than a million-token model if latency and cost are lower.

Technical details available from the release data include the 262,144-token context length, open-weight status, OpenRouter listing, and Ollama availability. Maximum output length, pricing, license specifics, and modality support are not specified here. Based on the listed capabilities, users should treat it primarily as a text/code model.

Laguna XS 2.1’s strength is focus. Poolside is associated with software engineering models, and a coding-oriented 262K context window is a useful middle ground: large enough for serious repository understanding, but potentially more efficient than ultra-long general models.

The limitation is that coding performance depends heavily on benchmarked correctness, tool use, test awareness, and edit quality—not just context size. Users should evaluate it on multi-file refactors, bug localization, test generation, and build-error repair before adopting it for autonomous coding workflows.

Kimi K2.7 Code: Moonshot’s long-context coding specialist

Kimi K2.7 Code is Moonshot AI’s coding-focused model with a 262,144-token context window. Like Laguna XS 2.1, it is available via OpenRouter and Ollama and is marked open weight.

Its listed capabilities are code generation, long-context reasoning, and software engineering-oriented understanding. That makes it relevant for tasks such as reading large repositories, explaining unfamiliar systems, generating patches across multiple files, or reasoning about dependency and API interactions in a codebase.

The main specs are a 262K-token context, open-weight availability, and code-focused capability profile. Pricing, maximum output, exact license terms, and modality support are not specified in the release data.

Kimi K2.7 Code’s benefit is specialization. Compared with general-purpose million-token models, a dedicated coding model may provide better completions, more idiomatic patches, and stronger reasoning about program structure—assuming Moonshot’s training and evaluation support those strengths.

Its caveats are familiar: code models can produce plausible but broken changes, miss hidden constraints, or overfit to visible patterns. Human review, test execution, and constrained editing workflows remain essential.

A practical note for software maintenance

Long-context and coding-focused models are increasingly useful for maintenance tasks such as dependency auditing, changelog summarization, version-drift detection, and repository-wide impact analysis. The best use is not to hand full control to a model, but to let it read more surrounding context, propose hypotheses, identify affected files, and explain likely upgrade risks. For production workflows, pair model output with deterministic tooling, tests, lockfile checks, and human review.

Bottom line

This week’s releases show long context becoming a standard competitive frontier rather than a niche feature. GLM-5.2 and Nemotron 3 Ultra stand out for combining open-weight access with million-token windows, Fugu Ultra adds another closed long-context reasoning option, and Laguna XS 2.1 plus Kimi K2.7 Code push repository-scale coding workflows forward.

The next phase will be less about who advertises the largest context window and more about who uses that context reliably, efficiently, and verifiably. Expect the strongest models to differentiate on long-context accuracy, tool integration, latency, cost control, and the ability to turn massive inputs into grounded, testable outputs.

Vibgrate CLI

See a real scan run

A replay of the actual CLI running against our test repositories — live progress, real findings, a genuine DriftScore. Nothing executes in your browser.

Replay
demo@vibgrate — bash
❯ npx @vibgrate/cli scan
 
╭──────────────────────────────────────────╮
│ Vibgrate Drift Report │
╰──────────────────────────────────────────╯
 
── node-turborepo (node) .
Runtime: >=18.0.0 (6 majors behind)
Frameworks:
Turbo: 1.13.4 → 2.11.5 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
1 current 1 1-behind 3 2+ behind 1 unknown
 
── @repo/admin (node) apps/admin
Frameworks:
TanStack Query: 5.104.0 → 5.104.0 (current)
React: 18.3.1 → 19.3.0 (1 behind)
React DOM: 18.3.1 → 19.3.0 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Vite: 5.4.21 → 8.3.1 (3 behind)
Dependencies:
3 current 9 1-behind 3 2+ behind 4 unknown
 
── @repo/api (node) apps/api
Frameworks:
Express: 4.22.3 → 5.2.1 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Vitest: 1.6.1 → 5.0.2 (4 behind)
Dependencies:
7 current 4 1-behind 4 2+ behind 4 unknown
 
── @repo/web (node) apps/web
Frameworks:
Next.js: 14.2.35 → 16.3.7 (2 behind)
React: 18.3.1 → 19.3.0 (1 behind)
React DOM: 18.3.1 → 19.3.0 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
2 current 6 1-behind 3 2+ behind 5 unknown
 
── @repo/config (node) packages/config
Frameworks:
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
2 current 2 1-behind 5 2+ behind 0 unknown
 
── @repo/database (node) packages/database
Frameworks:
Prisma: 5.22.0 → 7.10.0 (2 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
1 current 0 1-behind 3 2+ behind 1 unknown
 
── @repo/types (node) packages/types
Frameworks:
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Dependencies:
0 current 0 1-behind 1 2+ behind 1 unknown
 
── @repo/ui (node) packages/ui
Frameworks:
React: 18.3.1 → 19.3.0 (1 behind)
TypeScript: 5.9.3 → 7.0.2 (2 behind)
React: 18.3.1 → 19.3.0 (1 behind)
Dependencies:
1 current 4 1-behind 1 2+ behind 1 unknown
 
── @repo/utils (node) packages/utils
Frameworks:
TypeScript: 5.9.3 → 7.0.2 (2 behind)
Vitest: 1.6.1 → 5.0.2 (4 behind)
Dependencies:
0 current 1 1-behind 2 2+ behind 1 unknown
 
Tech Stack
Frontend: React, React DOM
Meta-frameworks: Next.js
Bundlers: tsx, Turbo, Vite
CSS / UI: Autoprefixer, PostCSS, Tailwind CSS
Backend: Express
ORM / Database: Prisma, Prisma Client
Testing: Vitest
Lint & Format: ESLint, ESLint Prettier, ESLint React, Prettier, typescript-eslint
 
Services & Integrations
Auth: JWT 9.0.3
Databases: Prisma 5.22.0
 
TypeScript
v5.3.3 · strict ✔ · MIXED · target: ES2022
 
Build & Deploy
Package Managers: pnpm
Monorepo: npm-workspaces, pnpm-workspaces, turbo
 
Product Purpose Signals
Frameworks: react, nextjs
Evidence: 177
Top Signals:
- [heading] Dashboard (apps/admin/src/pages/Dashboard.tsx)
- [title] Revenue Overview (apps/admin/src/pages/Dashboard.tsx)
- [copy] workspace:* (packages/ui/package.json)
- [copy] ./dist (packages/ui/tsconfig.json)
- [copy] ./src/index.ts (packages/ui/package.json)
- [copy] @repo/config/tsconfig-base.json (packages/ui/tsconfig.json)
- [copy] @repo/ui (packages/ui/package.json)
- [copy] #3b82f6 (apps/admin/src/pages/Dashboard.tsx)
Unknowns:
- No pricing or billing evidence found.
- No integrations/connectors evidence found.
- No route structure evidence found.
 
Security Posture
Lockfile ✖ · .env ✔ · node_modules ✔
 
Platform
Native modules: turbo
 
Code Quality
Files: 36 · Functions: 183 · Avg complexity: 2.62 · Avg length: 21.13 lines
Max nesting: 2 · Circular deps: 0 · Dead code: 0%
God files: apps/admin/src/pages/Products (448 lines)
 
Database Schema
postgresql · 8 models · 1 enum
Models: Address, CartItem, Category, Order, OrderItem (+3 more)
 
Findings (16 errors, 11 warnings)
✖ Node.js runtime ">=18.0.0" reached end-of-life on 2025-04-30 (latest: 24.0.0).
vibgrate/runtime-eol in .
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in .
✖ 60% of dependencies are 2+ major versions behind in node-turborepo.
vibgrate/dependency-rot in .
✖ @types/node is 6 major versions behind (spec: ^20.11.0, latest: 26.6.3).
vibgrate/dependency-major-lag in .
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in apps/admin
✖ Vite is 3 major versions behind (current: 5.4.21, latest: 8.3.1).
vibgrate/framework-major-lag in apps/admin
✖ vite is 3 major versions behind (spec: ^5.0.12, latest: 8.3.1).
vibgrate/dependency-major-lag in apps/admin
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in apps/api
✖ Vitest is 4 major versions behind (current: 1.6.1, latest: 5.0.2).
vibgrate/framework-major-lag in apps/api
✖ @types/node is 6 major versions behind (spec: ^20.11.0, latest: 26.6.3).
vibgrate/dependency-major-lag in apps/api
✖ vitest is 4 major versions behind (spec: ^1.2.1, latest: 5.0.2).
vibgrate/dependency-major-lag in apps/api
⚠ Next.js is 2 major versions behind (current: 14.2.35, latest: 16.3.7).
vibgrate/framework-major-lag in apps/web
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in apps/web
✖ @types/node is 6 major versions behind (spec: ^20.11.0, latest: 26.6.3).
vibgrate/dependency-major-lag in apps/web
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/config
✖ 56% of dependencies are 2+ major versions behind in @repo/config.
vibgrate/dependency-rot in packages/config
✖ eslint-plugin-react-hooks is 3 major versions behind (spec: ^4.6.0, latest: 7.1.1).
vibgrate/dependency-major-lag in packages/config
⚠ Prisma is 2 major versions behind (current: 5.22.0, latest: 7.10.0).
vibgrate/framework-major-lag in packages/database
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/database
✖ 75% of dependencies are 2+ major versions behind in @repo/database.
vibgrate/dependency-rot in packages/database
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/types
✖ 100% of dependencies are 2+ major versions behind in @repo/types.
vibgrate/dependency-rot in packages/types
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/ui
⚠ TypeScript is 2 major versions behind (current: 5.9.3, latest: 7.0.2).
vibgrate/framework-major-lag in packages/utils
✖ Vitest is 4 major versions behind (current: 1.6.1, latest: 5.0.2).
vibgrate/framework-major-lag in packages/utils
✖ 67% of dependencies are 2+ major versions behind in @repo/utils.
vibgrate/dependency-rot in packages/utils
✖ vitest is 4 major versions behind (spec: ^1.2.1, latest: 5.0.2).
vibgrate/dependency-major-lag in packages/utils
 
╭──────────────────────────────────────────╮
│ Top Priority Actions │
╰──────────────────────────────────────────╯
 
1. Upgrade EOL runtime in node-turborepo
End-of-life runtimes no longer receive security patches and block ecosystem upgrades.
./.
>=18.0.0 → 24.0.0 (6 majors behind)
Impact: −10 drift points (runtime & EOL)
 
2. Fix security posture: no lockfile found
Without a lockfile, installs are non-deterministic. Run the install command to generate one and commit it.
./
Missing: package-lock.json, pnpm-lock.yaml, or yarn.lock
 
3. Upgrade Vitest 1.6.1 → 5.0.2 in @repo/api (+2 more)
4 major versions behind. Major framework drift increases breaking change risk and blocks access to security fixes and performance improvements.
./apps/api
Vitest: 1.6.1 → 5.0.2 (4 majors behind)
./packages/utils
Vitest: 1.6.1 → 5.0.2 (4 majors behind)
./apps/admin
Vite: 5.4.21 → 8.3.1 (3 majors behind)
Impact: −5–15 drift points
 
4. Reduce dependency rot in @repo/types (100% severely outdated)
1 of 1 dependencies are 2+ majors behind. Run `npm outdated` and prioritise packages with known CVEs or breaking API changes.
./packages/types
typescript: 5.9.3 → 7.0.2 (2 majors behind)
Impact: −5–10 drift points
 
5. Reduce dependency rot in @repo/database (75% severely outdated)
3 of 4 dependencies are 2+ majors behind. Run `npm outdated` and prioritise packages with known CVEs or breaking API changes.
./packages/database
@prisma/client: 5.22.0 → 7.10.0 (2 majors behind)
prisma: 5.22.0 → 7.10.0 (2 majors behind)
typescript: 5.9.3 → 7.0.2 (2 majors behind)
Impact: −5–10 drift points
 
╭──────────────────────────────────────────╮
│ Architecture Layers │
╰──────────────────────────────────────────╯
 
Archetype: nextjs (80% confidence)
Files classified: 24 (11 unclassified)
Folders classified: 8
apps/admin/src presentation 100% 4 files
apps/admin/src/pages presentation 100% 2 files
apps/api/src/middleware middleware 100% 2 files
apps/api/src/routes routing 100% 2 files
apps/web/src/app presentation 100% 4 files
apps/web/src/app/products presentation 100% 2 files
apps/web/src/app/products/[id] presentation 100% 1 file
packages/ui/src presentation 100% 6 files
Unclassified source (sample): 11
 
presentation 15 files drift ████████████████████ 100 risk high
routing 4 files drift ████████████████████ 100 risk high
middleware 2 files drift ███████▍░░░░░░░░░░░░ 37 risk moderate
config 2 files drift ░░░░░░░░░░░░░░░░░░░░ 0 risk none
shared 1 file drift ████████████████████ 100 risk high
 
╭──────────────────────────────────────────╮
│ DriftScore Summary │
╰──────────────────────────────────────────╯
 
DriftScore: 70/100
Risk Level: HIGH
Projects: 9
Classified: 8 nano · 1 micro · 0 small · 0 standard
Billable: 0.42 · 9 detected → 0.42 billable projects (micro-project pricing)
0.1 micro · 0.32 nano
These fractions add up across repositories, then round down to whole billable projects.
 
Score Breakdown
Runtime: ████████████████████ 100
Frameworks: ███████████▊░░░░░░░░ 59
Dependencies: ██████▌░░░░░░░░░░░░░ 33
EOL Risk: ████████████████████ 100
 
Scanned at 2026-09-30T06:02:13.873Z · 7.2s · 286 files scanned · 56 workspace files · 27 dirs
❯
Press Run to start.