How Companies Operate Coding Agent Harnesses
Harness operations, team responsibilities, permissions, and verification across 12 companies, from Uber, BlackRock, Stripe, and Spotify to Microsoft, LG CNS, and Toss.
Keeping company standards consistent across coding tools requires separate controls for distributing instructions, enforcing permissions, and verifying repository changes.
Developers use Codex, Claude Code, and Cursor on laptops and remote machines. Even with identical instruction files, plugins, authentication, and local settings can change the tools available and the results.
This becomes harder across multiple teams. A payments team needs reconciliation rules, while a web team needs UI verification procedures. Common security policies must apply to both and remain in effect after tool updates.
This article examines configurations and responsibilities at 12 companies using official sources checked through October 2, 2026. The company diagrams reconstruct their published accounts; they are not internal deployment diagrams. The later sections on responsibilities and updates propose designs drawn from these cases.
What to manage centrally
An agent harness supplies context to the model and manages tool calls and task state. Organizations add installation, policies, internal knowledge, evaluation, and deployment around it. Supporting several products requires deciding which of these operations must remain consistent.
| Area | Standard to maintain | How to check |
|---|---|---|
| Instructions and business knowledge | Approved common standards and the team's knowledge | File versions and actual loading results |
| Skills and plugins | Reviewed procedures and executable code | Installation source, version, and required permissions |
| System access | Scope permitted for the user and task | Execute requests that should be allowed and denied |
| Result verification | Success criteria for each repository | Build, test, and review results |
| Updates | Existing workflows and required restrictions | Regression checks by product, OS, and execution environment |
Matching file contents addresses the first check. Whether a product reads the file, conflicting settings exist, or prohibited actions are blocked needs separate verification. The table does not promise identical model-generated code.
Public sources also establish different things. Engineering accounts describe operational configurations; vendor customer stories describe adoption as reported by the customer. Product documentation describes available capabilities and should be distinguished from evidence of a company's deployment.
Uber's shared execution environment and service ownership
The operating architecture published in August 2026 uses a common wrapper for each interactive coding harness. It manages installation, configuration, authentication, and usage visibility, while developers author task-specific skills. Uber reported more than 3,600 skills and 30,000 daily executions at publication.
Model selection uses internal benchmarks based on real PRs. The same operating framework tracks task results and the effects of model changes alongside shared installation. The figures describe Uber's reported usage, not productivity gains at other companies. Uber's development platform operations, 2026-08-27
The MCP architecture published in October separates the Registry and Gateway. The Registry manages tool definitions, owners, policies, and change history; the Gateway handles authentication, authorization, and responses. It centrally connects more than 5,000 tools, while service owners review their business meaning and whether to expose them.
Automatically discovered tools are registered as disabled. Description changes undergo owner review, deployment, and rollback procedures, and calls distinguish human, service, and agent identities. Platform operations are shared while responsibility for tools remains with each service. Uber MCP Gateway design, 2026-10-01
BlackRock's bread-kit and shared knowledge distribution
bread-kit provides development standards, architecture guidance, domain knowledge, and skills across development tools. It manages ownership and versions in Git, combining reviewed company knowledge with team and repository context. The management history of that knowledge persists when tools change.
In Cognition's customer story, the responsible BlackRock executive says all coding teams adopted it and use it to onboard new engineers. This is a customer account published by a vendor. It does not establish that BlackRock deployed every Devin Plugins management feature announced in the same article.
Exact user counts, IAM implementation, and version promotion procedures are undisclosed. The evidence establishes shared knowledge management in Git and distribution across tools. BlackRock bread-kit case, 2026-09-09
Stripe's Minions and defined execution procedures
Minions is an internal coding agent built on Goose. Its Blueprints connect agent decisions with deterministic steps such as linting and pushing. It reads Cursor rules scoped by directory and file pattern, and Stripe synchronizes those rules for Claude Code too.
Tasks receive a relevant subset of roughly 500 tools in the internal Toolshed MCP server. Prepared devboxes restrict access to production services, real user data, and unrestricted external networks. Local linting, CI, and bounded repair attempts precede returning results to a human. Stripe Minions Part 2, 2026-02-19
Requests start in existing collaboration workflows such as Slack and lead to review by the requesting engineer and another engineer. The amount of agent-written code and the scope of human review require separate interpretation. Stripe Minions Part 1
Spotify's Honk and repository-specific verification
Honk adds code transformation to Fleet Management/Fleetshift's existing repository selection, scheduling, PR creation, review, and merge workflow. Multiple entrypoints reuse the common CLI, logging, and tracing. The development platform that preceded agents continues to handle surrounding tasks. Honk Part 1, 2025-11-06
Repository contents determine which verifiers run, with different build and check procedures behind a common verifier interface. After those checks, a judge compares the request and changes; failed verification or a judge veto blocks PR creation. The article explicitly says the judge's own performance had not yet been evaluated. Honk Part 3, 2025-12-09
The follow-up architecture published in June 2026 uses the Claude Agent SDK, Spotify's harness, and Kubernetes. It connects tasks to Backstage component ownership, documentation, and existing development standards. The limitations of the 2025 implementation and the 2026 architecture describe different points in time. Spotify's development environment for teams and agents, 2026-06-03
Block's company and module checks
Block provides a common sq agents review entrypoint on developer machines and in the cloud. It runs module owners' checks in .agents/checks/*.md alongside company checks. Nested AGENTS.md files supply code context, and an internal Skills Marketplace shares procedures.
Each check runs in parallel through a review agent with independent context, and the results are combined. Incidents and announcements can prompt policy and check proposals, but humans review changes and give final code approval. The platform handles shared execution and result collection; teams familiar with the code manage check contents. Block's check and skill operations, 2026-04-02
Meta's performance optimization team and domain skills
Meta's published case concerns performance optimization in its Capacity Efficiency organization. MCP tools provide profiling, experiment results, configuration history, and code search. Skills hold expertise for specific codebases, languages, and problem types. Optimization and regression response reuse the same tools.
Optimization candidates are checked for syntax, style, and fit to the problem, then presented in the engineer's editor. Regression-fix PRs request review from the engineer who authored the causal change. The source does not describe automatic merging, production deployment, or company-wide IAM implementation.
This case shows how a specialist organization builds workflows on common tools. It does not establish identical instruction distribution, permission inheritance, or version promotion across all Meta development teams. Meta Capacity Efficiency case, 2026-04-16
Google's code migrations and code-owner review
Google published an internal tool built by Core Developer, Ads, and DeepMind, including an Ads migration of IDs from 32 to 64 bits. Experts identify targets using Code Search, Kythe, and scripts. The tool expands the set of related files, and Gemini, fine-tuned on internal code, generates candidate edits.
Multiple prompts generate candidates in parallel, and compilation and unit tests evaluate them to select changes. The model attempts repairs when needed. Engineers inspect and edit the results, then split changes for review by affected code owners. Initial target discovery and expansion did not use AI at that time.
Domain experts, repository verification, and review remain part of large migrations. This source describes migration work at the time. It does not establish current general-purpose coding harnesses or company-wide MCP Gateway operations. Google Research code migrations, 2024-07-18
Microsoft's risk-based management and central support
Microsoft Digital's internal account distinguishes personal, team, and enterprise-service agents. It assesses data access together with actions: low-risk reads can use self-service, while changes to organizational systems require more development and security review.
Team agents need ownership and lifecycle management. This is internal agent governance broader than coding harnesses. Microsoft internal governance, 2026-05-21
The AI CoE began with training and advice, then expanded into priorities, common standards, and execution coordination. Microsoft explicitly says it uses Agent 365 internally to view agent inventories, owners, lifecycles, and governance status across platforms. Microsoft AI CoE, 2026-04-16
Its MCP account covers server registration, ownership, review, and short-lived, scoped tokens. Gateway placement is also described as a recommendation, end-to-end logging as a development goal, and automatic blocking of unregistered servers as an ideal state. These are not proof that all unapproved calls are already blocked and audited company-wide. Microsoft MCP security and governance
| Company | Shared operating unit | Team or task responsibility | Scope of the source |
|---|---|---|---|
| Uber | Wrapper, Registry, Gateway | Service tool review and skill authoring | Internal developer platform |
| BlackRock | Git-managed bread-kit | Team and repository knowledge | Customer account published by a vendor |
| Stripe | Blueprint, Toolshed, devbox | Task context and code review | Internal Minions tasks |
| Spotify | Fleet workflows, shared execution and tracing | Repository verification criteria | Honk and later development environments |
| Block | Shared check execution and skill distribution | Module checks and approval | Developer and cloud review workflows |
| Meta | Common tools and data access | Performance optimization domain skills | Capacity Efficiency organization |
| Migration generation and evaluation procedures | Expert scope selection and code review | Migrations published in 2024 | |
| Microsoft | Risk-based policies and central support | Agent ownership and lifecycle | Internal governance broader than coding |
LG CNS's development knowledge and financial project
Agentic AIND, published in June 2026, organizes development standards, security rules, source code, and artifacts into a Knowledge Foundation. Analysis, design, coding, and testing agents divide responsibilities. They implement and verify defined specifications, with human review and approval.
LG CNS says it is applying COBOL-to-Java conversion capabilities to a next-generation project at a large Korean financial company. This does not mean deployment to every customer. The source does not establish team policy inheritance, MCP permissions, or version promotion implementation. LG CNS announcement, 2026-06-08
A separate February announcement reported a joint study of 26 real projects. Average productivity gains were 26.1% for the earlier DevOn AIND and 14.1% for general-purpose tools. These are company-reported figures; the original paper's detailed experimental design was not verified here. They must not be treated as effects of the full Agentic AIND feature set released in June. LG CNS research announcement, 2026-02-12
Samsung SDS's coding assistance and AI Crews
Samsung SDS says it uses IDE-integrated AI Pro and CodeBot for coding assistance and reviews. In 2025, it assigned at least one AI Crew member to each team alongside a central organization. This connected local use-case discovery and adoption with governance.
The same source describes agents connecting planning, generation, review, verification, and system changes as being prepared or built. MCP and internal permission integration were also under development. It publishes pilot results but says operational adoption needs further security and CX review. Tool use and adoption structures are documented, but a completed company-wide agent control architecture is not established. Samsung SDS AI Journey, 2025-10-24
NAVER's team assets and Context Provider
NAVER D2's June 2026 source presents the Place AI recommendation organization's experience building a Context Provider. The platform automatically collects team assets from data and serving layers to supply context to people and AI agents.
The verified scope covers the presentation description, the team's affiliation, and its application topic. This is not an analysis of the video's detailed implementation. It cannot establish MCP, IAM, or deployment architecture, or company-wide coding harness operations at NAVER. NAVER D2 presentation, 2026-06-22
Toss's experiment sharing and MCP for external developers
AI Surf Day shared experiments through 142 Evangelists across subsidiaries and teams and about 200 Clubs at the start. These are people and organization counts, not harness deployments. A Codex workflow combined implementation, builds, iOS Simulator interaction, edits, and video inspection. It was a hackathon winner, not evidence of the entire daily development process. Toss AI Surf Day, 2026-06-05
Payments MCP is a knowledge tool that helps external merchant developers integrate payments. It follows document versions through the existing MDX-to-CDN CI/CD pipeline and distributes a local STDIO server through npm. This documents management of both source knowledge and its delivery path. Payments MCP implementation, 2025-06-24
The two sources cover internal experiment sharing and external developer support respectively. They do not establish a harness applying the same skills, policies, and permissions to every internal Toss development team.
| Company | Published scope | What the source alone does not establish |
|---|---|---|
| LG CNS | Structured development knowledge, role-specific tasks, application to a financial project | Team policy inheritance, MCP permissions, version promotion implementation |
| Samsung SDS | Coding assistance tools and team adoption structure | A completed company-wide coding-agent control architecture |
| NAVER | One team's experience collecting assets and building a Context Provider | Deployment and permission management of a company-wide coding harness |
| Toss | AI experiment sharing across the group and Payments MCP for external developers | A common harness operating across all internal development teams |
Dividing responsibility when applying the cases
The following sections propose designs drawn from these cases. Companies published different scopes at different times; no single company should be assumed to implement the entire structure. Shared operations and ownership of business knowledge can be divided as follows.
| Scope | Owner | Responsibilities |
|---|---|---|
| Company | Developer platform team | Supported tools, installation and deployment, common execution and evaluation tools |
| Security and identity | Security, IAM, and data owners | Access policies, credentials, required restrictions, exception procedures |
| Domain | Service and domain owners | Business terminology, API contracts, domain skills |
| Team and repository | Team and code owners | Local instructions, workflows, builds, tests, and checks |
| Individual | Developer | Permitted preferences such as language and output format |
Selecting a team's skill must not expand access permissions. Installing a reconciliation procedure, for example, must leave production ledger writes subject to a separate policy decision. Adding more specific knowledge and changing permissions require separate controls.
The following example generalizes that division of responsibility. It shows model connections separately from MCP tool calls.
Registering a server in the catalog does not grant permission to call it. Successful Gateway authentication still requires checking the resources permitted by the connected system. MCP security guidance prohibits accepting and forwarding tokens that were not issued for the target server. MCP security recommendations
Local files, shells, browsers, and model communication also have their own execution paths. Assuming that MCP Gateway logs capture every agent action leaves gaps in observability.
Why product settings need verification
Applying common policies across products requires checking what each setting means. Ordinary defaults, model instructions, and restrictions that users cannot change are different controls. The table describes differences in official product documentation, not customer deployment outcomes.
| Product | Management behavior to check | Difference that is easy to miss |
|---|---|---|
| Codex | Roles of requirements.toml and ordinary config.toml | The command network proxy does not cover all MCP, app, and web-search communication |
| Claude Code | Sources of managed settings and merge behavior | Default source precedence differs from the separate merge option |
| GitHub Copilot | Supported clients for each setting | Support differs across CLI, IDE, app, and cloud |
| Cursor | How organization group settings combine | Several settings adopt the more permissive value |
Sources are Codex managed configuration, Claude settings deployment, Copilot managed settings support, and Cursor organization groups.
A hook error does not always stop execution. Claude and Copilot may continue after some hook timeouts or communication failures. Hooks used for permission controls need tests for failure behavior as well as successful responses. Claude hooks, Copilot hooks
OS differences also matter. Claude's Bash sandbox does not support native Windows. Consider restricting unsupported combinations to an execution environment that can enforce the required controls. Claude sandbox
Product adapters need both configuration translation and tests of the resulting behavior. Explicitly marking unsupported requirements reduces the risk of mistaking a substitute setting for the requested control. After version changes, verify that the same setting keys still produce the same behavior.
Separating team packages from permission changes
Team-managed packages can record their scope, owner, required tools, and verification methods. The example shows management fields, not an actual Codex or Claude configuration schema.
id: payments-reconciliation
owner: payments-platform
version: 1.4.0
scope:
repositories: [payments-ledger]
context:
instructions: instructions.md
skills: [reconcile-in-test]
requires:
tools: [ledger-test-read, reconciliation-test-run]
verification:
checks: [ledger-invariants, reconciliation-integration]requires declares the tools needed; it does not grant permissions. Access must depend on authenticated user and task identities and service policies. Changing a team name in a local file must not expand access.
| Change | Reviewer | What to verify |
|---|---|---|
| Team documentation and code explanations | Team and repository owners | Accuracy and applicable location |
| Repeated-task skills | Team owner | Results and scope on representative tasks |
| Executable scripts and hooks | Platform owner, with security review when needed | Execution permissions and error and timeout handling |
| MCP tools and descriptions | Service owner | Input and output contracts and exposed data |
| Expanded production writes or external transfers | Security, data, and service owners | Approved scope, auditing, and revocation |
Central approval for every sentence edit slows improvement. Changes containing executable code or permission updates need review by the owners of affected systems. Routing reviews by change type preserves team autonomy alongside common controls.
Evaluating updates against actual tasks
Evaluation needs a more specific target than a product name. Record the client version, OS, execution environment, model, common policies, team package, and tool contracts together. Treat a product's CLI and app, or native Windows and WSL2, as separate combinations.
Uber's uReview uses human-labeled evaluation data and expanded incrementally by team and review function. It also incorporates user feedback into evaluation. That case can inform an update procedure such as the following. Uber uReview
Configuration-file inspection alone cannot establish actual behavior. Test successful permitted tasks alongside rejection of prohibited tasks. The following regression checks should be adapted to the organization's requirements.
| Check | Question |
|---|---|
| Common instruction loading | Did the approved version enter the session? |
| Permitted tasks | Can the agent complete representative tasks using the required documents and tools? |
| Prohibited tasks | Are unauthorized file and system access denied? |
| Authentication or policy server outage | Do required restrictions hold, or does work stop in the specified way? |
| Hook errors and timeouts | Does failure behavior match the designed policy? |
| Existing sessions | Was immediate policy application or the need to restart verified? |
| Result quality | Are review effort, corrections, rework, and defects increasing? |
Pinning an old version cannot maintain service compatibility indefinitely. Codex app update controls and Cursor deployment policies also have limits on their scope and supported versions. Operations need a way to validate updates and a procedure for returning to the previous approved state. Codex update management, Cursor deployment management
Installation rates and generated-code volume alone do not establish performance. Observe review time, rework, change failures, and recovery too. DORA's analysis also examines time saved during generation shifting into a verification burden. DORA's AI adoption analysis
Summary
The published cases connect team and repository knowledge and checks to shared execution and deployment infrastructure. Applying this approach requires separate checks for instruction loading, actual permission decisions, and result verification. Product-specific configuration precedence and failure behavior also belong in the operating contract.
A small rollout can start with one team's representative task and complete verification. Adding another team means connecting its business knowledge and checks while verifying that common controls still hold. Updates can then expand from combinations that pass the same task and prohibited-action checks.