Ecosystem
July 22, 2026

AWS shares reference architecture for offline-first generative AI at the edge

AWS has published a reference architecture for building offline-first generative AI applications that combine cloud-based model customization with local edge inference for reliable, low-latency AI in disconnected environments.

AWS has introduced a reference architecture for deploying offline-first generative AI applications across edge environments with unreliable or intermittent connectivity.

The approach combines Amazon Bedrock for training data generation, Amazon SageMaker AI for model fine tuning, AWS IoT Greengrass for deployment orchestration, and local inference using Ollama and Strands Agents.

A hybrid fine tuning plus retrieval augmented generation (RAG) strategy enables domain-specific responses while keeping models compact enough for edge hardware. The architecture also includes continuous feedback loops, cloud-to-edge synchronization, and security controls for authentication, encryption, prompt guardrails, and monitoring across manufacturing, energy, agriculture, and other remote operations.

#
AWS

Read Our Content

See All Blogs
AI safety

Enterprise AI security: How GoML builds prompt injection-resistant applications

Paushigaa S

July 21, 2026
Read more
Gen AI

How Hiswai built an AI report generation software with GoML to cut research time by 80%

Deveshi Dabbawala

July 20, 2026
Read more