Cover photo by Matheus Bertelli on Pexels.
For the past two years, teams in Latin America have lived with an awkward compromise: the most capable open models came with licenses that lawyers had to read twice. You could experiment freely, but the moment a pilot started handling real customer messages or real revenue, someone had to ask whether the terms allowed it — and under what conditions. That question shaped architecture decisions, vendor choices, and even whether a project stayed on a local server or moved to a third-party API.
Gemma 4 changes the shape of that conversation. Announced on April 2, 2026, Google DeepMind's new open model family moves to Apache 2.0 — a commercially permissive license — and spans efficient on-device builds to server-grade models, with 140+ language coverage across the family. The license is the headline, but the range of sizes is what makes it usable: you can prototype on a large model and then check whether a smaller one handles your narrow task with retrieval and guardrails.
Why the license matters
Earlier Gemma releases used Gemma Terms, which carried use restrictions that mattered most exactly when a project turned commercial. Gemma 4 moves to Apache 2.0. In practical terms, you can integrate, modify, and commercialize under that license without a separate commercial agreement with Google, subject to the license text and the model terms.
Read those documents before production use — the official Gemma 4 announcement, the DeepMind model page, and the Hugging Face collection all link the current terms. And keep in mind what a license does not do: it does not remove your duties around privacy, security, and local regulation, including Brazil's LGPD. If your workflow touches customer data, run compliance review in parallel with technical evaluation, not after it.
The current model family
The April 2 launch announced four variants under Apache 2.0. The current Google model card now lists five, including 12B Unified:
| Variant | Context | Practical distinction |
|---|
| E2B | 128K | Small efficient model. |
| E4B | 128K | Larger efficient model. |
| 12B Unified | 256K | Dense 12B-class model with audio support. |
| 26B A4B | 256K | Mixture of experts; about 3.8B parameters active. |
| 31B | 256K | Largest dense model in this family. |
Memory use also depends on precision, quantization and serving software. Measure it on the exact build you intend to deploy; parameter count alone is not a hardware recommendation.

Choose the build that fits the workflow, not the other way around. Photo by Andrea Piacquadio on Pexels.
What Google’s benchmarks actually say
| Benchmark | 31B | 26B A4B |
|---|
| MMLU Pro | 85.2 | 82.6 |
| AIME 2026, no tools | 89.2 | 88.3 |
| LiveCodeBench v6 | 80.0 | 77.1 |
| GPQA Diamond | 84.3 | 82.3 |
| MMMLU | 88.4 | 86.3 |
| MMMU Pro, multimodal | 76.9 | 73.8 |
| Tau2, average over three domains | 76.9 | 68.2 |
These are author-reported scores in the Google model card, not our tests or a prediction of business results. MMLU Pro and MMMLU are different evaluations; the old article incorrectly used 85.2 for MMMLU. Ongoing leaderboard positions are not fixed product specifications. The April 2 launch announcement reported Arena positions using its April 1 snapshot: third among open models for 31B and sixth for 26B. These are dated launch positions, not current rankings or evaluations run by TakeAICourse. The table above uses the current Google model card, checked October 6, 2026. MMMU Pro measures multimodal understanding; Tau2 averages three tool-use domains. Keep these evaluations separate when selecting a pilot.
What is practical in Latin America
The useful combination is permissive licensing plus multilingual coverage plus a range from efficient to server models. Three patterns are worth testing:
- Spanish and Brazilian Portuguese assistants where you keep data and prompts under your control. Local or VPC deployment keeps sensitive conversations inside infrastructure you govern — relevant wherever third-party APIs are difficult to approve.
- Regulated workflows with review gates. Drafting, classification, and summarization tasks where every output passes a human or rule-based check before it reaches a customer.
- Cost tiering. Prototype on 31B or 26B-A4B, then check whether E4B or E2B handles the narrow task with retrieval and guardrails. Treat smaller-model adequacy as a testable hypothesis for your task.
Scale will depend on local evaluation, fine-tunes, Vertex AI / Hugging Face / edge runtime support, and compliance review — not the license alone. If your team serves both Spanish- and Portuguese-speaking users, test both: multilingual coverage is a starting point, and tone, formality, and domain vocabulary still need checking per variant.
Try it without rebuilding your stack
Start from the official Hugging Face Gemma 4 collection, the Google developer announcement (Apr 2, 2026), and the DeepMind model page and the current model card. Pick one workflow — for example, drafting replies to ten real anonymized support messages in your variant of Spanish or Portuguese — then score accuracy, tone, hallucinations, and time saved.

Score every trial the same way: accuracy, tone, hallucinations, time saved. Photo by Daniil Komov on Pexels.
Record the model version, prompt, temperature, and dataset so a colleague can repeat the test. That log matters more than the first result: it turns a demo into an evaluation you can rerun when the next model version arrives. If the small model passes with retrieval, you have your deployment answer. If it does not, you know exactly which cases need the larger build — or a different approach entirely.
If you need structure for that evaluation habit, continue with the guide hub and turn this story into one reusable prompt in the prompt library rather than browsing at random. For guided practice with a current product, see Claude Code: Ship Your First App or the current course catalog. You can buy one paid course for US$20 or choose school access. Gemma hosting and compute are separate costs.