The numbers say 768p. MiniMax H3, the video generation model the market calls open source, tops out at 768p on local hardware. The 2K module exists. It is confirmed. It produces high-resolution video. But it sits behind an API, reachable only on MiniMax's servers. That gap is not a bug report. It is an architecture statement.
I have spent 23 years reading deployment order as evidence. When a team says "we are open-sourcing the model" and then gates the premium capability behind a managed service, the sequencing is the data. A 2K frame at 2560x1440 holds roughly 3.7 million pixels. A 768p frame at 1366x768 holds about 1.05 million. That is a 3.5x pixel gap. For a video transformer, where self-attention cost grows with token count, the inference gap is closer to an order of magnitude. The math does not weep, it merely liquidates. The liquidation order here: quality at the server, cost at the edge.
Let me be precise about what we actually know. The source is a Reddit AMA, forwarded through external media monitoring. Self-reported team communication. Not a technical paper. Not a third-party evaluation. A first-phase breakdown yields nine information points, concentrated in product roadmap, feature availability, and known defects. Four facts are confirmed. H3 generates complete 768p video locally. A 2K video module is confirmed but only accessible via the official API. A local acceleration solution is planned. And the team admits multimodal joint referencing and long-shot small-figure scenes show blurring and distortion. Missing entirely: model architecture, parameter count, training data, evaluation metrics, pricing, release timeline, license type. When an AI company discloses a 2K API but omits its license, the omission is itself a data point.
I filter this the way I audit smart contracts. Rank the relevance dimensions. High: technical roadmap, commercialization, infrastructure and compute. Medium: competitive landscape, industry impact. Low: investment narratives and ethics — the source says nothing on those, so this analysis does not speculate. Confidence is C-level. Reasonable inference, pending disclosure.
The hidden information matters more than the declared facts. A 2K module that re-processes rather than generates implies the high-resolution capability is a separate model family. That is why it can ship as an API without retraining the base model. Compute cost is why it must. The choice of redraw over simple upscaling signals semantic reconstruction — a higher ceiling for text and face restoration, but a higher risk of hallucination and content drift. If redraw changes what is on screen, it is not a lossless enlarge. It is a re-creation. For frame-consistency-critical work, that is a new liability.
The core evidence structure breaks into four findings. First, the architecture is two-tier. H3 generates end-to-end at 768p. The 2K module does not generate 2K from text or images. It re-processes existing video and original reference material to restore text, faces, and scene detail. That is a post-processing lane. The 2K module is architecturally independent of the base model. You do not retrain a 768p foundation model to ship a 2K endpoint. You train a separate refinement model, deploy it on your own hardware, and meter access.
Second, the compute economics confirm the rollout order. Native 2K generation on consumer hardware is economically prohibitive. API-first is the correct engineering call. Scaling makes it worse: attention is computed across spatial tokens per frame and temporal tokens across frames, so the 2K path multiplies both the spatial field and the conditioning context. Memory bandwidth becomes the wall. The team's promised local acceleration is not optional polish; it is the prerequisite for any local 2K workflow. The sequence is the map — 2K API first, local acceleration second, complete local 2K workflow "someday." Each step reveals a cost constraint. If 2K redraw were cheap, the API gate would not exist.
Third, the business model is open-core, whether they call it that or not. Base capability open. Premium capability gated. This is the playbook I have watched in crypto for a decade: open base layer, toll booth at the value layer. The local acceleration plan is not charity; it is user acquisition. Developers build on local 768p, integrate H3 into pipelines, and when a client demands 2K, the official API is the only exit ramp. The headline says open source. The functional reality: the best output of the model runs on hardware the user does not control. Open weights plus a closed premium tier is a business model, not a contribution to open research. I do not predict the future, I verify the past — and every open-core product in crypto eventually reaches a license cliff.
Fourth, the defects are model-level, not pipeline-level. The team attributed blurring and deformation to the model itself. That places the failure in multimodal condition encoding and spatial-temporal generation, not in rendering or post-processing. This connects to the redraw route. The 2K module is not interpolation. It is semantic reconstruction. Every frame is a re-interpretation, and a semantic redraw can silently alter identity between cuts. The AMA discloses no temporal consistency numbers. After my 2017 audit season — 15 contracts, 42 critical vulnerabilities — I learned that what teams omit from their own reports is often the first thing to fail.
But the open questions outnumber the confirmed facts. Does the 2K redraw preserve temporal consistency across cuts? Is the 2K module larger than the 768p base model? What is the acceleration path — distillation, quantization, pruning, caching, or sparse temporal attention? What frame rate, VRAM footprint, and generation latency remain after acceleration? And is 768p itself truly end-to-end, or does it hide its own multi-stage pipeline? None of these have answers. In a bull market for AI video, that is the part that gets skipped.

Now kill two assumptions. First, do not conflate open source with decentralized. Running 768p weights on your own GPU is not verified inference. A video model is stochastic. Its output cannot be proven on-chain the way a transaction can. In 2026, I designed a zero-knowledge protocol to verify AI-generated data trails — one million model outputs processed — and it worked only because the data was deterministic. Video generation is not. Verification of provenance is not verification of quality. If a redraw model synthesizes a face from twenty pixels, no oracle can tell you which face is true. For anyone building AI-crypto rails, this is the blind spot. The euphoria does not change the architecture. Every FOMO buyer of the story should read the AMA the way I read a vesting contract: check what is unlocked, check what is locked, and check who owns the master key.
Second, the paywall may be temporary. API-first correlates with inference cost, not ideology. If decentralized GPU networks mature — and the supply is growing — or if distillation, pruning, and sparse temporal attention cut 2K cost by an order of magnitude, the local 2K module becomes a competitive weapon and the weights open. The open-core verdict is C-level confidence, not A. I have been burned assuming a vendor's constraint is permanent. Compute is a variable. Variables get optimized.
The near-term industry impact lands where quality tolerance is low: short-form video, ad creative, concept previews. The acknowledged distortion keeps this out of film-grade and commercial-delivery pipelines. If the 2K redraw validates, it will penetrate subtitle restoration, old-footage repair, and e-commerce detail video — workflows that already accept post-processing as standard. And note where the real margin sits: data-sensitive clients in film, advertising, and regulated industries cannot push footage into a third-party API. Private deployment converts 2K access into a commercial negotiation. The local acceleration scheme is an enterprise play, not a consumer courtesy.
Here is the signal. Six months. If local acceleration ships and the 2K weights remain behind the API, open-core is confirmed, and the pattern is set for every AI-video model that follows. Until then, the chain of custody is incomplete — no license, no parameter count, no temporal consistency metric, no pricing. Treat the H3 release as a status message, not a deliverable. The 2K gate will not be lifted by goodwill. It will be lifted by a cheaper compute curve, an open competitor, or a decentralized GPU supply that undercuts the API. One of those will arrive. When the best output of an open model requires a server the user does not control, openness is not a property of the weights. It is a state of access. Liquidity is not a promise, it is a state of flow — and so is open source. Watch the API pricing page, not the announcement thread. The past does not change its verdict. The toll booth always arrives.