English

PLUS ULTRAProduct LaunchesStrandsStrands harness

Strands Releases Open-Source Agent Harness for Efficient Local and Cloud Deployment

PLUS ULTRA by Amenoyomi

Strands has released Strands harness, an open-source, fully assembled agent harness designed for both local execution and cloud deployment. Available under an Apache 2.0 license, the harness allows developers to implement state-of-the-art agents using a single line of Python or TypeScript code.

The harness is built to be a general-purpose tool rather than being limited to coding tasks. It supports a wide range of model providers, including Amazon Bedrock, Anthropic, OpenAI, Google, and local models via Ollama. For deployment, it is compatible with various Linux container services such as Modal, Cloudflare Containers, Azure Container Apps, Google Cloud Run, and Amazon ECS.

In benchmarking tests, Strands harness demonstrated significant cost efficiency. According to the company, it was 28% cheaper than other harnesses when using the same Claude or GPT models across six benchmarks. When compared specifically to Claude Code using Fable 5, Strands harness was 77% more cost-effective and achieved higher scores on the Terminal Bench 2.1.

The framework includes built-in defaults for prompt caching and context management to maintain accuracy while optimizing token usage. Its context management features include truncating tool results over ~1500 tokens, triggering summarization when the context window exceeds 85%, and performing context recovery during overflows.

For developers who require deeper control, the Strands Harness SDK provides granular management over the agent loop, tools, memory, and multi-agent patterns. Additionally, the Strands CLI allows for rapid prototyping in plain English, with an /export command to convert CLI prototypes into Python or TypeScript code.

PLUS ULTRAby Amenoyomi

The reduction in operational costs is primarily driven by the harness's default context management system, which optimizes how tokens are consumed during agent loops. Rather than relying solely on model selection, the system employs three specific automated controls to limit unnecessary token growth.

First, tool results exceeding approximately 1,500 tokens are automatically truncated to prevent oversized inputs from bloating the context. Second, a summarization process, referred to as compaction, is triggered once the context window reaches 85% capacity. Finally, if a context overflow occurs, a recovery mechanism runs within the loop to maintain stability.

This approach differs from other high-efficiency harnesses, such as the DeepSeek Harness, which the company notes is the most token-efficient overall but typically reports the lowest accuracy scores. By using these specific management defaults, Strands harness maintains equal or better accuracy across benchmarks while reducing the cost of using models like Claude and GPT.

Sources

  1. Strands Harness (Hacker News Frontpage, 2026-09-23)
  2. Strands Harness SDK