PROJECT
OSVA
Open infrastructure for AI agents, workflows, and AI workforces.
What is it?
EXECUTION
Container exited
The runtime process was gone
OUTPUT
Artifact still readable
Its bytes remained available
AUTHORITY
Metadata in PostgreSQL
The execution record survived
The test proved that an artifact could outlive the container that created it. This is test evidence, not a public demo.
OSVA is an open-source platform for building, running, controlling, observing, and evaluating AI agents and agentic workflows. It is an agent operating layer, not only a framework: versioned agents, governed model access, workflows with branching, parallel steps and approval gates, cron scheduling, MCP client and server, and knowledge retrieval. The OSS 1.0 source release is Apache 2.0 and self-hosts with Docker Compose or Helm.
Here is the point I care about most, and it is a small one. The agent's container can exit. The artifact it produced is still available.
In our container artifact integration test, an agent created an artifact through OSVA and got back a reference. After the container exited, the bytes were still readable. The authoritative record of that artifact lives in PostgreSQL, not in the process that made it.
To be straight about it, that is verified by an integration test today, not by a click-through demo. The demo is still on my list.
The thesis behind the whole project fits in one sentence. Agent execution should be a managed product capability, with explicit state, permissions, and recovery.
Source →Why we built it?
An agent finishes a task on a developer's machine. Good. Now turn that into a product and a different set of questions shows up. Where does its output live? What can it access? What happens when execution fails halfway? How does a workflow pick back up after waiting on something outside the system?
Generating an answer and operating an agent are two different responsibilities. Most demos only do the first one.
I want to be careful about what I am claiming. I am not saying existing agent frameworks lack orchestration or persistence. Plenty of them are good at it. The boundary I chose is different. OSVA owns execution infrastructure and the public contracts around it, and agent logic can use its own libraries behind the runtime interface.
Three bets sit underneath the design. They are judgments, and I could be wrong about any of them.
Developers will need freedom to change how agents run. So runtime execution is separated from the platform's durable state. A particular agent implementation should not quietly become the architecture.
Recovery and permissions will matter as soon as an agent does anything useful. I would rather build them into the platform boundary than leave every application to reconstruct them on its own.
An open-source edition has to solve a complete problem. The Community Edition should be useful by itself. Organizational controls and deployment requirements are reserved for future commercial editions.
I built OSVA as an open-source product foundation. It also gives me a concrete way to talk about product scope, architecture, security, and build-versus-buy decisions in leadership and advisory conversations.
How it works?
Seven concepts hold the whole thing together. Once these click, the rest of the page makes sense.
Agent. A definition of the work an agent can perform and how it should run.
Run. A tracked execution of an agent against an input.
Attempt. An individual execution attempt within a run.
Runtime. The environment that executes agent code, such as a container or an HTTP service.
Workflow. A defined process connecting work, including waiting for external events.
Capability. An operation the runtime can request through OSVA without ever receiving the underlying credentials.
Artifact. A stored output with durable metadata and a reference that other work can use.
Durable authority
PostgreSQL holds the authoritative durable state. BullMQ and Valkey provide job transport.
The technology names matter less than the reasoning. A queue message should not be the only record of what work exists or what happened to it. And process memory should not decide whether a run or a workflow survives.
Here is what one run looks like from submission to recorded outcome.
Execution separated from credentials
Agent runtimes do not receive database, vector-store, provider, or secret credentials. They ask OSVA to perform supported operations through capability interfaces.
That draws an explicit line between executing agent code and touching platform resources.
I want to be clear about the limit here. This is a boundary in the design, not a claim of complete sandbox security. Runtime isolation and deployment policy still matter.
Waiting is a durable state
A workflow can enter WAITING, persist that wait, and resume when a matching event arrives. Event handling is idempotent, which means the same event delivered twice should not cause the same effect twice.
The product reasoning is simple. Waiting on something outside the system is part of the process, so it needs a durable representation, not a sleeping thread.
The operating system analogy
The metaphor holds in three places: execution management, durable state, and permission boundaries. I am not going to force a one-to-one mapping with a computer operating system, because past those three places it stops being true.
How to use it?
Setup
cd deploy/compose
cp .env.example .env
# Set POSTGRES_PASSWORD in .env
docker compose up -d --build --wait
docker compose run --rm migrate
docker compose --profile bootstrap run --rm bootstrap
export OSVA_API_KEY='osva_ak_…' # from bootstrap output
curl -sS -H "Authorization: Bearer $OSVA_API_KEY" \
http://127.0.0.1:8080/v1/api-keys | jq .
cd /path/to/osva
pnpm install
pnpm exec turbo run build --filter=@osva/sdk
export OSVA_BASE_URL=http://127.0.0.1:8080
export OSVA_API_KEY='osva_ak_…'
node --input-type=module -e "
import { OsvaClient } from '@osva/sdk';
const client = new OsvaClient({ baseUrl: process.env.OSVA_BASE_URL, apiKey: process.env.OSVA_API_KEY });
const keys = await client.apiKeys.list();
console.log(keys);
"
python -m pip install ./sdks/python
export OSVA_BASE_URL=http://127.0.0.1:8080
export OSVA_API_KEY='osva_ak_…'
python -c "
import os
from osva import OSVAClient
client = OSVAClient(base_url=os.environ['OSVA_BASE_URL'], api_key=os.environ['OSVA_API_KEY'])
print(client.agents.list())
"The demo I am planning to build
A product UI, a demo video, and the reproducible six-step flow below are not built yet. The repository includes developer-facing runtime samples, but not the onboarding demo I intend to ship. Until then, this is the demo I am planning, and it is labeled that way on purpose.
Start OSVA
Launch OSVA using the documented local setup.
Register the sample agent
Add a bounded container agent that produces a file as its output, such as converting structured input into a JSON report.
Submit a run
Send demo input through the SDK or API and receive a Run identifier.
Watch execution
See the Run and Attempt move through dispatch and container execution via API responses, CLI output, or logs.
Create the artifact
The agent requests artifact creation through the OSVA capability boundary and returns an ArtifactReference.
Retrieve it after exit
The container stops, then retrieve the artifact through OSVA and confirm its bytes and metadata remain available.
The demo will make the infrastructure promise visible: execution is temporary, durable state and outputs are not.
The last two steps are the whole point. They land the hook: the container is gone and the artifact is still there.
Using OSVA on a real project
- Choose one bounded task. That way success and failure are both obvious.
- Define the agent's input and output contract. Callers should know what to expect.
- Select a runtime that fits your deployment. Container or HTTP service, depending on the requirements.
- Configure required capabilities. Do not pass infrastructure credentials into agent code.
- Store reusable outputs as artifacts. Their lifecycle should be independent of the runtime.
- Use workflows where work spans steps or external events.
- Test failure and repeated delivery before you depend on the process.
- Choose storage and deployment settings that suit the environment. Filesystem locally, S3 for object storage.
A few things I would tell you before you start
Reliable infrastructure does not guarantee useful model output. OSVA can run an agent dependably and the agent can still give a bad answer.
Container networking needs an explicit operator policy. Persisted artifacts bring storage and lifecycle responsibilities with them. External operations need their own idempotency strategy. And passing local acceptance checks proves specific behavior, not production scale.
What important decisions we took while building?
Six decisions shaped the design.
| Decision | Reason | Cost or tradeoff |
|---|---|---|
| PostgreSQL as durable authority | Keep state independent of queues and running processes. | Database migrations and persistence discipline. |
| BullMQ and Valkey for transport | Separate dispatch from authoritative business state. | Coordination between transport and durable records. |
| Capability-mediated access | Keep underlying credentials outside agent runtimes. | More explicit interfaces and integration work. |
| Stable public contracts | Give SDKs and runtimes a consistent boundary. | Versioning and compatibility obligations. |
| Filesystem and S3 artifact storage | Support local use and object-storage deployments. | Two adapters, plus integrity and streaming concerns. |
| A useful Community Edition | Make the open-source product independently valuable. | A more careful commercial boundary. |
The test that failed, and what it taught me.
BEFORE
Execution existed only in memory
In-memory Agent
↓
In-memory Run
↓
In-memory Attempt
↓
Artifact capability request
↓
Generic artifact capability failure
DURABILITY ASSUMED, NOT PROVEN
Artifact metadata referenced durable execution records that did not exist in PostgreSQL.
Persist the complete execution context
AFTER
Execution records persisted first
Agent in PostgreSQL
↓
Run in PostgreSQL
↓
Attempt in PostgreSQL
↓
Container creates Artifact
↓
ArtifactReference returned
↓
Container exits
↓
Artifact bytes remain readable
END-TO-END OUTCOME VERIFIED
The failure showed that testing artifact creation alone was not enough. The product outcome depended on the full execution record being durable.
In the container artifact integration test, the agent reached the artifact capability, but artifact creation failed with a generic error. Nothing pointed at the real cause.
The cause was in the test itself. It had represented its execution records only in memory. But artifact metadata has foreign-key relationships that require the agent, the run, and the attempt to exist in PostgreSQL. In memory, they did not exist where the database could see them.
I fixed the test setup to persist those records. The integration then passed: the container created an artifact, returned its reference, and the bytes stayed readable after the container exited.
What stayed with me is the product lesson. An artifact belongs to a tracked execution. Testing the capability in isolation was not enough to prove the outcome a user actually cares about.
Release checks should give the same answer on every machine. A migration-history check hashed raw files, and the hashes changed with Windows line endings. We moved to canonical LF hashing so the same migration produces the same hash everywhere. Small fix, but a release check that disagrees with itself is not a check.
What we deliberately didn't build
Every item here was a choice. The cut list says as much about the project as the feature list does.
Infrastructure credentials inside runtimes. It would be easier to hand an agent a database connection. I chose capability requests instead, because that keeps the boundary explicit.
Provider-specific types in the domain. Provider details stay outside the core contracts so the platform does not become tied to one provider.
Process memory as durable authority. Anything that must survive a restart lives in persisted state. No exceptions for convenience.
SDKs that bypass platform boundaries. SDKs talk over HTTP and public contracts, the same way any other client would.
Enterprise controls in the Community scope. SSO, SCIM, advanced organizational controls, and enterprise deployment requirements belong to the commercial roadmap.
Known limitations
Owned, not hidden.
What has been checked. Version 1.0.0 across the recorded release components. 33 of 33 implementation slices completed. Quick verification, clean CI verification, migration verification, release-readiness checks, and Linux kind acceptance all passing.
What that does not prove. It is engineering acceptance evidence. It does not establish adoption, production scale, or package publication.
| Limitation | Impact | Next step |
|---|---|---|
| UI work and testing remained unfinished | Infrastructure evidence exceeds end-user experience evidence. | Build and test the main user journeys. |
| Guided onboarding examples and sample workflows were missing | New users must bridge documentation and practical use themselves. | Ship runnable, bounded examples. |
| No confirmed public demo | Visitors cannot immediately experience the project. | Publish a reproducible demo. |
| Acceptance testing is not production operating history | Passing checks does not establish sustained reliability or scale. | Gather operational and load evidence. |
| Agent quality remains application-dependent | Reliable execution can still produce poor decisions or answers. | Evaluate each use case and define review boundaries. |
The honest bottleneck: OSVA can make execution more dependable. It cannot decide whether an agent's output is good enough for a particular business decision. That part stays human.
This is the kind of product work I want to lead: turning technical capability into a system people can use, inspect, and trust.