Add GPTSwarm (Graph-based Workflow) (#2460 )

* update * revert the main branch lock file * regenerate the poetry.lock * move poetry into another dependency group * fix infer.sh and gpt code --------- Co-authored-by: yufansong <yufan@risingwave-labs.com>
2026-04-29 03:00:45 -04:00 · 2024-08-10 23:02:44 -07:00
2992 changed files with 104477 additions and 373498 deletions
--- a/.agents/skills/cross-repo-testing/SKILL.md
+++ b/.agents/skills/cross-repo-testing/SKILL.md
@@ -1,202 +0,0 @@
---
-name: cross-repo-testing
-description: This skill should be used when the user asks to "test a cross-repo feature", "deploy a feature branch to staging", "test SDK against OH Cloud", "e2e test a cloud workspace feature", "test provider tokens", "test secrets inheritance", or when changes span the SDK and OpenHands server repos and need end-to-end validation against a staging deployment.
-triggers:
- cross-repo
- staging deployment
- feature branch deploy
- test against cloud
- e2e cloud
---
-
-# Cross-Repo Testing: SDK ↔ OpenHands Cloud
-
-How to end-to-end test features that span `OpenHands/software-agent-sdk` and `OpenHands/OpenHands` (the Cloud backend).
-
-## Repository Map
-
-| Repo | Role | What lives here |
-|------|------|-----------------|
-| [`software-agent-sdk`](https://github.com/OpenHands/software-agent-sdk) | Agent core | `openhands-sdk`, `openhands-workspace`, `openhands-tools` packages. `OpenHandsCloudWorkspace` lives here. |
-| [`OpenHands`](https://github.com/OpenHands/OpenHands) | Cloud backend | FastAPI server (`openhands/app_server/`), sandbox management, auth, enterprise integrations. Deployed as OH Cloud. |
-| [`deploy`](https://github.com/OpenHands/deploy) | Infrastructure | Helm charts + GitHub Actions that build the enterprise Docker image and deploy to staging/production. |
-
-**Data flow:** SDK client → OH Cloud API (`/api/v1/...`) → sandbox agent-server (inside runtime container)
-
-## When You Need This
-
-There are **two flows** depending on which direction the dependency goes:
-
-| Flow | When | Example |
-|------|------|---------|
-| **A — SDK client → new Cloud API** | The SDK calls an API that doesn't exist yet on production | `workspace.get_llm()` calling `GET /api/v1/users/me?expose_secrets=true` |
-| **B — OH server → new SDK code** | The Cloud server needs unreleased SDK packages or a new agent-server image | Server consumes a new tool, agent behavior, or workspace method from the SDK |
-
-Flow A only requires deploying the server PR. Flow B requires pinning the SDK to an unreleased commit in the server PR **and** using the SDK PR's agent-server image. Both flows may apply simultaneously.
-
---
-
-## Flow A: SDK Client Tests Against New Cloud API
-
-Use this when the SDK calls an endpoint that only exists on the server PR branch.
-
-### A1. Write and test the server-side changes
-
-In the `OpenHands` repo, implement the new API endpoint(s). Run unit tests:
-
-```bash
-cd OpenHands
-poetry run pytest tests/unit/app_server/test_<relevant>.py -v
-```
-
-Push a PR. Wait for the **"Push Enterprise Image" (Docker) CI job** to succeed — this builds `ghcr.io/openhands/enterprise-server:sha-<COMMIT>`.
-
-### A2. Write the SDK-side changes
-
-In `software-agent-sdk`, implement the client code (e.g., new methods on `OpenHandsCloudWorkspace`). Run SDK unit tests:
-
-```bash
-cd software-agent-sdk
-pip install -e openhands-sdk -e openhands-workspace
-pytest tests/ -v
-```
-
-Push a PR. SDK CI is independent — it doesn't need the server changes to pass unit tests.
-
-### A3. Deploy the server PR to staging
-
-See [Deploying to a Staging Feature Environment](#deploying-to-a-staging-feature-environment) below.
-
-### A4. Run the SDK e2e test against staging
-
-See [Running E2E Tests Against Staging](#running-e2e-tests-against-staging) below.
-
---
-
-## Flow B: OH Server Needs Unreleased SDK Code
-
-Use this when the Cloud server depends on SDK changes that haven't been released to PyPI yet. The server's runtime containers run the `agent-server` image built from the SDK repo, so the server PR must be configured to use the SDK PR's image and packages.
-
-### B1. Get the SDK PR merged (or identify the commit)
-
-The SDK PR must have CI pass so its agent-server Docker image is built. The image is tagged with the **merge-commit SHA** from GitHub Actions — NOT the head-commit SHA shown in the PR.
-
-Find the correct image tag:
- Check the SDK PR description for an `AGENT_SERVER_IMAGES` section
- Or check the "Consolidate Build Information" CI job for `"short_sha": "<tag>"`
-
-### B2. Pin SDK packages to the commit in the OpenHands PR
-
-In the `OpenHands` repo PR, pin all 3 SDK packages (`openhands-sdk`, `openhands-agent-server`, `openhands-tools`) to the unreleased commit and update the agent-server image tag. This involves editing 3 files and regenerating 3 lock files.
-
-Follow the **`update-sdk` skill** → "Development: Pin SDK to an Unreleased Commit" section for the full procedure and file-by-file instructions.
-
-### B3. Wait for the OpenHands enterprise image to build
-
-Push the pinned changes. The OpenHands CI will build a new enterprise Docker image (`ghcr.io/openhands/enterprise-server:sha-<OH_COMMIT>`) that bundles the unreleased SDK. Wait for the "Push Enterprise Image" job to succeed.
-
-### B4. Deploy and test
-
-Follow [Deploying to a Staging Feature Environment](#deploying-to-a-staging-feature-environment) using the new OpenHands commit SHA.
-
-### B5. Before merging: remove the pin
-
-**CI guard:** `check-package-versions.yml` blocks merge to `main` if `[tool.poetry.dependencies]` contains `rev` fields. Before the OpenHands PR can merge, the SDK PR must be merged and released to PyPI, then the pin must be replaced with the released version number.
-
---
-
-## Deploying to a Staging Feature Environment
-
-The `deploy` repo creates preview environments from OpenHands PRs.
-
-**Option A — GitHub Actions UI (preferred):**
-Go to `OpenHands/deploy` → Actions → "Create OpenHands preview PR" → enter the OpenHands PR number. This creates a branch `ohpr-<PR>-<random>` and opens a deploy PR.
-
-**Option B — Update an existing feature branch:**
-```bash
-cd deploy
-git checkout ohpr-<PR>-<random>
-# In .github/workflows/deploy.yaml, update BOTH:
-#   OPENHANDS_SHA: "<full-40-char-commit>"
-#   OPENHANDS_RUNTIME_IMAGE_TAG: "<same-commit>-nikolaik"
-git commit -am "Update OPENHANDS_SHA to <commit>" && git push
-```
-
-**Before updating the SHA**, verify the enterprise Docker image exists:
-```bash
-gh api repos/OpenHands/OpenHands/actions/runs \
-  --jq '.workflow_runs[] | select(.head_sha=="<COMMIT>") | "\(.name): \(.conclusion)"' \
-  | grep Docker
-# Must show: "Docker: success"
-```
-
-The deploy CI auto-triggers and creates the environment at:
-```
-https://ohpr-<PR>-<random>.staging.all-hands.dev
-```
-
-**Wait for it to be live:**
-```bash
-curl -s -o /dev/null -w "%{http_code}" https://ohpr-<PR>-<random>.staging.all-hands.dev/api/v1/health
-# 401 = server is up (auth required). DNS may take 1-2 min on first deploy.
-```
-
-## Running E2E Tests Against Staging
-
-**Critical: Feature deployments have their own Keycloak instance.** API keys from `app.all-hands.dev` or `$OPENHANDS_API_KEY` will NOT work. You need a test API key issued by the specific feature deployment's Keycloak.
-
-**You (the agent) cannot obtain this key yourself** — the feature environment requires interactive browser login with credentials you do not have. You must **ask the user** to:
-1. Log in to the feature deployment at `https://ohpr-<PR>-<random>.staging.all-hands.dev` in their browser
-2. Generate a test API key from the UI
-3. Provide the key to you so you can proceed with e2e testing
-
-Do **not** attempt to log in via the browser or guess credentials. Wait for the user to supply the key before running any e2e tests.
-
-```python
-from openhands.workspace import OpenHandsCloudWorkspace
-
-STAGING = "https://ohpr-<PR>-<random>.staging.all-hands.dev"
-
-with OpenHandsCloudWorkspace(
-    cloud_api_url=STAGING,
-    cloud_api_key="<test-api-key-for-this-deployment>",
-) as workspace:
-    # Test the new feature
-    llm = workspace.get_llm()
-    secrets = workspace.get_secrets()
-    print(f"LLM: {llm.model}, secrets: {list(secrets.keys())}")
-```
-
-Or run an example script:
-```bash
-OPENHANDS_CLOUD_API_KEY="<key>" \
-OPENHANDS_CLOUD_API_URL="https://ohpr-<PR>-<random>.staging.all-hands.dev" \
-python examples/02_remote_agent_server/10_cloud_workspace_saas_credentials.py
-```
-
-### Recording results
-
-Both repos support a `.pr/` directory for temporary PR artifacts (design docs, test logs, scripts). These files are automatically removed when the PR is approved — see `.github/workflows/pr-artifacts.yml` and the "PR-Specific Artifacts" section in each repo's `AGENTS.md`.
-
-Push test output to the `.pr/logs/` directory of whichever repo you're working in:
-```bash
-mkdir -p .pr/logs
-python test_script.py 2>&1 | tee .pr/logs/<test_name>.log
-git add -f .pr/logs/
-git commit -m "docs: add e2e test results" && git push
-```
-
-Comment on **both PRs** with pass/fail summary and link to logs.
-
-## Key Gotchas
-
-| Gotcha | Details |
-|--------|---------|
-| **Feature env auth is isolated** | Each `ohpr-*` deployment has its own Keycloak. Production API keys don't work. Agents cannot log in — you must ask the user to provide a test API key from the feature deployment's UI. |
-| **Two SHAs in deploy.yaml** | `OPENHANDS_SHA` and `OPENHANDS_RUNTIME_IMAGE_TAG` must both be updated. The runtime tag is `<sha>-nikolaik`. |
-| **Enterprise image must exist** | The Docker CI job on the OpenHands PR must succeed before you can deploy. If it hasn't run, push an empty commit to trigger it. |
-| **DNS propagation** | First deployment of a new branch takes 1-2 min for DNS. Subsequent deploys are instant. |
-| **Merge-commit SHA ≠ head SHA** | SDK CI tags Docker images with GitHub Actions' merge-commit SHA, not the PR head SHA. Check the SDK PR description or CI logs for the correct tag. |
-| **SDK pin blocks merge** | `check-package-versions.yml` prevents merging an OpenHands PR that has `rev` fields in `[tool.poetry.dependencies]`. The SDK must be released to PyPI first. |
-| **Flow A: stock agent-server is fine** | When only the Cloud API changes, `OpenHandsCloudWorkspace` talks to the Cloud server, not the agent-server. No custom image needed. |
-| **Flow B: agent-server image is required** | When the server needs new SDK code inside runtime containers, you must pin to the SDK PR's agent-server image. |
--- a/.agents/skills/custom-codereview-guide.md
+++ b/.agents/skills/custom-codereview-guide.md
@@ -1,47 +0,0 @@
---
-name: custom-codereview-guide
-description: Repo-specific code review guidelines for All-Hands-AI/OpenHands. Provides frontend and backend review rules in addition to the default code review skill.
-triggers:
- /codereview
---
-
-# All-Hands-AI/OpenHands Code Review Guidelines
-
-You are an expert code reviewer for the **All-Hands-AI/OpenHands** repository. This skill provides repo-specific review guidelines.
-
-## Frontend: i18n / Translation Key Usage
-
-**Never dynamically construct i18n keys via string interpolation or template literals.**
-
-All translation keys must come from the `I18nKey` enum (`frontend/src/i18n/declaration.ts`) or from canonical mapping objects like `AGENT_STATUS_MAP` (`frontend/src/utils/status.ts`). Dynamically constructed keys (e.g., `` t(`STATUS$${value.toUpperCase()}`) ``) will silently fall back to the raw key string at runtime because `i18next` returns the key itself when a translation is missing — this produces broken UI text with no build-time or test-time error.
-
-### What to flag
-
- Any call to `t(...)` or `i18next.t(...)` where the key is built at runtime via template literals, string concatenation, or helper functions rather than referencing `I18nKey` or a known mapping
- Any new i18n key referenced in code that does not exist in `frontend/src/i18n/translation.json`
-
-### Correct pattern
-
-```ts
-import { AGENT_STATUS_MAP } from "#/utils/status";
-
-const i18nKey = AGENT_STATUS_MAP[agentState];
-const message = i18nKey ? t(i18nKey) : fallback;
-```
-
-### Incorrect pattern
-
-```ts
-// BAD: constructs a key that may not exist in translation.json
-const message = t(`STATUS$${agentState.toUpperCase()}`);
-```
-
-## Frontend: Data Fetching Architecture
-
-UI components must never call API client methods (`frontend/src/api/`) directly. All data access must go through TanStack Query hooks:
-
-```
-UI components → TanStack Query hooks (frontend/src/hooks/query/ or mutation/) → API client (frontend/src/api/) → API endpoints
-```
-
-Flag any component that imports directly from `#/api/` and calls fetch/mutation functions without a TanStack Query wrapper.
--- a/.agents/skills/upcoming-release/SKILL.md
+++ b/.agents/skills/upcoming-release/SKILL.md
@@ -1,37 +0,0 @@
---
-name: upcoming-release
-description: This skill should be used when the user asks to "generate release notes", "list upcoming release PRs", "summarize upcoming release", "/upcoming-release", or needs to know what changes are part of an upcoming release.
---
-
-# Upcoming Release Summary
-
-Generate a concise summary of PRs included in the upcoming release.
-
-## Prerequisites
-
-Two commit SHAs are required:
- **First SHA**: The older commit (current release)
- **Second SHA**: The newer commit (what's being released)
-
-If the user does not provide both SHAs, ask for them before proceeding.
-
-## Workflow
-
-1. Run the script from the repository root with the `--json` flag:
-   ```bash
-   .github/scripts/find_prs_between_commits.py <older-sha> <newer-sha> --json
-   ```
-
-2. Filter out PRs that are:
-   - Chores
-   - Dependency updates
-   - Adding logs
-   - Refactors
-
-3. Categorize the remaining PRs:
-   - **Features** - New functionality
-   - **Bug fixes** - Corrections to existing behavior
-   - **Security/CVE fixes** - Security-related changes
-   - **Other** - Everything else
-
-4. Format the output with PRs listed under their category, including the PR number and a brief description.
--- a/.agents/skills/update-sdk/SKILL.md
+++ b/.agents/skills/update-sdk/SKILL.md
@@ -1,123 +0,0 @@
---
-name: update-sdk
-description: This skill should be used when the user asks to "update SDK", "bump SDK version", "pin SDK to a commit", "test unreleased SDK", "update agent-server image", "bump the version", "prepare a release", "what files change for a release", or needs to know how SDK packages are managed in the OpenHands repository. For detailed reference material, see references/docker-image-locations.md and references/sdk-pinning-examples.md in this skill directory.
---
-
-# Update SDK
-
-Bump SDK packages (`openhands-sdk`, `openhands-agent-server`, `openhands-tools`), pin them to unreleased commits for testing, and cut an OpenHands release.
-
-## Quick Summary — How Many Files Change?
-
-| Activity | Manual edits | Auto-regenerated | Total |
-|----------|:------------:|:----------------:|:-----:|
-| **SDK bump** (released PyPI version) | 2 | 3 | **5** |
-| **SDK pin** (unreleased git commit) | 3 | 3 | **6** |
-| **Release commit** (version bump) | 3 | 0 | **3** |
-
-The 3 auto-regenerated files are always: `poetry.lock`, `uv.lock`, `enterprise/poetry.lock`.
-
-## SDK Package Bump — 2 Files + 3 Lock Files
-
-Land as a separate PR before the release. Examples: `929dcc3` (SDK 1.11.5), `cd235cc` (SDK 1.11.4).
-
-| File | What to change |
-|------|----------------|
-| `pyproject.toml` | `openhands-sdk`, `openhands-agent-server`, `openhands-tools` in **two** sections: the `dependencies` array (PEP 508) **and** `[tool.poetry.dependencies]` |
-| `openhands/app_server/sandbox/sandbox_spec_service.py` | `AGENT_SERVER_IMAGE` constant — set to `ghcr.io/openhands/agent-server:<version>-python` |
-
-Then regenerate lock files:
-```bash
-poetry lock && uv lock && cd enterprise && poetry lock && cd ..
-```
-
-## Docker Image Locations — All Hardcoded References
-
-For the complete inventory of every file containing a hardcoded Docker image tag or repository, see `references/docker-image-locations.md`. Key files that must stay in sync during an SDK bump:
-
-| File | Image reference | Updated during SDK bump? |
-|------|----------------|:------------------------:|
-| `openhands/app_server/sandbox/sandbox_spec_service.py` | `AGENT_SERVER_IMAGE = 'ghcr.io/openhands/agent-server:<tag>-python'` | ✅ Yes |
-| `docker-compose.yml` | `AGENT_SERVER_IMAGE_TAG` default | ✅ Should be |
-| `containers/dev/compose.yml` | `AGENT_SERVER_IMAGE_REPOSITORY` + `_TAG` defaults | ✅ Should be |
-
-> **CI enforcement:** `.github/workflows/check-version-consistency.yml` validates version consistency and compose file image references on every PR and push to main.
-
-### ⚠️ Docker Image Tag Gotcha (merge-commit SHA)
-
-The SDK CI in `software-agent-sdk` repo tags Docker images with the **GitHub Actions merge-commit SHA**, NOT the PR head-commit SHA. When pinning to an SDK PR branch:
-
-1. Check the SDK PR description for the actual image tag (look for the `AGENT_SERVER_IMAGES` section)
-2. Or query the CI logs: the "Consolidate Build Information" job prints `"short_sha": "<tag>"`
-3. The merge-commit SHA differs from the head SHA shown in the PR
-
-For released SDK versions, images use a version tag (e.g., `1.12.0-python`) — no merge-commit ambiguity.
-
-## Cutting a Release — 3 Files
-
-A release commit updates the version string across 3 files. Gold-standard examples: 1.3.0 (`d063c8c`), 1.4.0 (`495f48b`).
-
-| File | What to change |
-|------|----------------|
-| `pyproject.toml` | `version = "X.Y.Z"` under `[tool.poetry]` |
-| `frontend/package.json` | `"version": "X.Y.Z"` |
-| `frontend/package-lock.json` | `"version": "X.Y.Z"` in **two** places (root object and `packages[""]`) |
-
-> **Note:** `openhands/version.py` reads the version from `pyproject.toml` at runtime — no manual edit needed there.
-
-### Compose Files (2 files)
-
-Both compose files should use `ghcr.io/openhands/agent-server` with the current SDK version tag.
-
-| File | What to verify |
-|------|----------------|
-| `docker-compose.yml` | `AGENT_SERVER_IMAGE_REPOSITORY` defaults to agent-server, `AGENT_SERVER_IMAGE_TAG` is current |
-| `containers/dev/compose.yml` | Same — must use agent-server, not runtime |
-
-### Release Workflow
-
-#### Step 1: Verify the SDK bump has landed
-
-```bash
-grep -n "openhands-sdk\|openhands-agent-server\|openhands-tools" pyproject.toml
-grep -n "AGENT_SERVER_IMAGE" openhands/app_server/sandbox/sandbox_spec_service.py
-grep "AGENT_SERVER_IMAGE_TAG" docker-compose.yml containers/dev/compose.yml
-```
-
-#### Step 2: Bump version numbers
-
-```bash
-# Edit pyproject.toml, frontend/package.json, frontend/package-lock.json
-git add pyproject.toml frontend/package.json frontend/package-lock.json
-git commit -m "Release X.Y.Z"
-git tag X.Y.Z
-```
-
-Create a `saas-rel-X.Y.Z` branch from the tagged commit for the SaaS deployment pipeline.
-
-#### Step 3: Images get tagged automatically
-
-Every push to `main` / `saas-rel-*` / `oss-rel-*` builds and publishes `ghcr.io/openhands/openhands` and `ghcr.io/openhands/enterprise-server` images for that commit (tagged by SHA, short SHA, and branch name).
-
-Pushing a git tag `X.Y.Z` then tags the images for that commit with `X.Y.Z`, `X.Y`, `X`, and `latest`. Non-semver tags just get their literal name applied.
-
-Requires the commit to already be built. If you push the tag too early, the retag CI job fails loudly — re-run it from the Actions UI once the build completes.
-
-## Development: Pin SDK to an Unreleased Commit
-
-For detailed examples of all pinning formats (commit, branch, uv-only), see `references/sdk-pinning-examples.md`.
-
-### Files to change (3 manual + 3 lock files)
-
-| File | What to change |
-|------|----------------|
-| `pyproject.toml` | Pin all 3 SDK packages in **both** `dependencies` and `[tool.poetry.dependencies]` |
-| `openhands/app_server/sandbox/sandbox_spec_service.py` | `AGENT_SERVER_IMAGE` — use the merge-commit SHA tag, NOT the head-commit SHA |
-| `docker-compose.yml` | `AGENT_SERVER_IMAGE_TAG` default (for local development) |
-| `poetry.lock` | Auto-regenerated via `poetry lock` |
-| `uv.lock` | Auto-regenerated via `uv lock` |
-| `enterprise/poetry.lock` | Auto-regenerated via `cd enterprise && poetry lock` |
-
-### CI guard
-
-The `check-package-versions.yml` workflow blocks merging to `main` if `[tool.poetry.dependencies]` contains any `rev` fields. This ensures unreleased SDK pins do not accidentally ship in a release.
--- a/.agents/skills/update-sdk/references/docker-image-locations.md
+++ b/.agents/skills/update-sdk/references/docker-image-locations.md
@@ -1,61 +0,0 @@
-# Docker Image Locations — Complete Inventory
-
-Every file in the OpenHands repository containing a hardcoded Docker image tag, repository, or version-pinned image reference. Organized by update cadence.
-
-## Updated During SDK Bump (must change)
-
-These files contain image tags that **must** be updated whenever the SDK version or pinned commit changes.
-
-### `openhands/app_server/sandbox/sandbox_spec_service.py`
- **Line:** `AGENT_SERVER_IMAGE = 'ghcr.io/openhands/agent-server:<tag>-python'`
- **Format:** `<sdk-version>-python` for releases (e.g., `1.12.0-python`), `<7-char-commit-hash>-python` for dev pins
- **Source of truth** for which agent-server image the app server pulls at runtime
- **⚠️ Gotcha:** When pinning to an SDK PR, the image tag is the **merge-commit SHA** from GitHub Actions, not the PR head-commit SHA. Check the SDK PR description or CI logs for the correct tag.
-
-### `docker-compose.yml`
- **Lines:**
-  ```yaml
-  - AGENT_SERVER_IMAGE_REPOSITORY=${AGENT_SERVER_IMAGE_REPOSITORY:-ghcr.io/openhands/agent-server}
-  - AGENT_SERVER_IMAGE_TAG=${AGENT_SERVER_IMAGE_TAG:-<tag>-python}
-  ```
- Used by `docker compose up` for local development
-
-### `containers/dev/compose.yml`
- **Lines:**
-  ```yaml
-  - AGENT_SERVER_IMAGE_REPOSITORY=${AGENT_SERVER_IMAGE_REPOSITORY:-ghcr.io/openhands/agent-server}
-  - AGENT_SERVER_IMAGE_TAG=${AGENT_SERVER_IMAGE_TAG:-<tag>-python}
-  ```
- Used by the dev container setup
- **Known issue:** On main as of 1.4.0, this file still points to `ghcr.io/openhands/runtime` instead of `agent-server`, and the tag is `1.2-nikolaik` (stale from the V0 era). The `check-version-consistency.yml` CI workflow catches this.
-
-## Updated During Release Commit (version string only)
-
-### `pyproject.toml`
- **Line:** `version = "X.Y.Z"` under `[tool.poetry]`
- The Python version is derived from this at runtime via `openhands/version.py`
-
-### `frontend/package.json`
- **Line:** `"version": "X.Y.Z"`
-
-### `frontend/package-lock.json`
- **Two places:** root `"version": "X.Y.Z"` and `packages[""].version`
-
-## Dynamic References (auto-derived, no manual update)
-
-### `openhands/version.py`
- Reads version from `pyproject.toml` at runtime → `openhands.__version__`
-
-### `.github/scripts/update_pr_description.sh`
- Uses `${SHORT_SHA}` variable at CI runtime, not hardcoded
-
-### `enterprise/Dockerfile`
- `ARG BASE="ghcr.io/openhands/openhands"` — base image, version supplied at build time
-
-## Image Registries
-
-| Registry | Usage |
-|----------|-------|
-| `ghcr.io/openhands/agent-server` | V1 agent-server (sandbox) — built by SDK repo CI |
-| `ghcr.io/openhands/openhands` | Main app image — built by `ghcr-build.yml` |
-| `docker.openhands.dev/openhands/*` | Mirror/CDN for the above images |
--- a/.agents/skills/update-sdk/references/sdk-pinning-examples.md
+++ b/.agents/skills/update-sdk/references/sdk-pinning-examples.md
@@ -1,103 +0,0 @@
-# SDK Pinning Examples
-
-Examples from real commits showing how to pin SDK packages to unreleased commits, branches, or released versions.
-
-## Pin to a Specific Commit
-
-Example from commit `169fb76` (pinning all 3 packages to SDK commit `100e9af`):
-
-### `dependencies` array (PEP 508 format)
-
-```toml
-"openhands-agent-server @ git+https://github.com/OpenHands/software-agent-sdk.git@100e9af#subdirectory=openhands-agent-server",
-"openhands-sdk @ git+https://github.com/OpenHands/software-agent-sdk.git@100e9af#subdirectory=openhands-sdk",
-"openhands-tools @ git+https://github.com/OpenHands/software-agent-sdk.git@100e9af#subdirectory=openhands-tools",
-```
-
-### `[tool.poetry.dependencies]` (Poetry format)
-
-```toml
-openhands-sdk = { git = "https://github.com/OpenHands/software-agent-sdk.git", rev = "100e9af", subdirectory = "openhands-sdk" }
-openhands-agent-server = { git = "https://github.com/OpenHands/software-agent-sdk.git", rev = "100e9af", subdirectory = "openhands-agent-server" }
-openhands-tools = { git = "https://github.com/OpenHands/software-agent-sdk.git", rev = "100e9af", subdirectory = "openhands-tools" }
-```
-
-### `openhands/app_server/sandbox/sandbox_spec_service.py`
-
-```python
-AGENT_SERVER_IMAGE = 'ghcr.io/openhands/agent-server:<merge-commit-sha>-python'
-```
-
-**⚠️ Important:** The image tag is the **merge-commit SHA** from the SDK CI, not the commit hash used in `pyproject.toml`. Look up the correct tag from the SDK PR description or CI logs.
-
-## Pin to a Branch
-
-Example from commit `430ee1c` (pinning to branch `openhands/issue-2228-sdk-settings-schema`):
-
-### `[tool.poetry.dependencies]`
-
-```toml
-openhands-sdk = { git = "https://github.com/OpenHands/software-agent-sdk.git", branch = "openhands/issue-2228-sdk-settings-schema", subdirectory = "openhands-sdk" }
-openhands-agent-server = { git = "https://github.com/OpenHands/software-agent-sdk.git", branch = "openhands/issue-2228-sdk-settings-schema", subdirectory = "openhands-agent-server" }
-openhands-tools = { git = "https://github.com/OpenHands/software-agent-sdk.git", branch = "openhands/issue-2228-sdk-settings-schema", subdirectory = "openhands-tools" }
-```
-
-## Using `[tool.uv.sources]` Override
-
-When only `uv` needs the override (keep PyPI versions in the main arrays), add a `[tool.uv.sources]` section. Example from commit `1daca49`:
-
-```toml
-[tool.uv.sources]
-openhands-sdk = { git = "https://github.com/OpenHands/software-agent-sdk.git", subdirectory = "openhands-sdk", rev = "4170cca" }
-openhands-agent-server = { git = "https://github.com/OpenHands/software-agent-sdk.git", subdirectory = "openhands-agent-server", rev = "4170cca" }
-openhands-tools = { git = "https://github.com/OpenHands/software-agent-sdk.git", subdirectory = "openhands-tools", rev = "4170cca" }
-```
-
-## Released PyPI Version (standard release)
-
-Example from commit `929dcc3` (SDK 1.11.5):
-
-### `dependencies` array
-
-```toml
-"openhands-agent-server==1.11.5",
-"openhands-sdk==1.11.5",
-"openhands-tools==1.11.5",
-```
-
-### `[tool.poetry.dependencies]`
-
-```toml
-openhands-sdk = "1.11.5"
-openhands-agent-server = "1.11.5"
-openhands-tools = "1.11.5"
-```
-
-### `openhands/app_server/sandbox/sandbox_spec_service.py`
-
-For released versions, the image tag uses the version number:
-
-```python
-AGENT_SERVER_IMAGE = 'ghcr.io/openhands/agent-server:1.11.5-python'
-```
-
-However, **some releases use a commit-hash tag** even for the released version. Check which tag format exists on GHCR. Example from `929dcc3`:
-
-```python
-AGENT_SERVER_IMAGE = 'ghcr.io/openhands/agent-server:010e847-python'
-```
-
-## Regenerate Lock Files
-
-After any change to `pyproject.toml`, always regenerate:
-
-```bash
-poetry lock
-uv lock
-cd enterprise && poetry lock && cd ..
-```
-
-## CI Guards
-
- **`check-package-versions.yml`**: Blocks merge to `main` if `[tool.poetry.dependencies]` contains `rev` fields (prevents shipping unreleased SDK pins)
- **`check-version-consistency.yml`**: Validates version strings match across `pyproject.toml`, `package.json`, `package-lock.json`, and verifies compose files use `agent-server` images
--- a/.devcontainer/README.md
+++ b/.devcontainer/README.md
@@ -1 +0,0 @@
-This way of running OpenHands is not officially supported. It is maintained by the community.
--- a/.devcontainer/devcontainer.json
+++ b/.devcontainer/devcontainer.json
@@ -1,19 +0,0 @@
-// For format details, see: https://aka.ms/devcontainer.json
-{
-	"name": "Python 3",
-	// Documentation for this image:
-	// - https://github.com/devcontainers/templates/tree/main/src/python
-	// - https://github.com/microsoft/vscode-remote-try-python
-	// - https://hub.docker.com/r/microsoft/devcontainers-python
-	"image": "mcr.microsoft.com/devcontainers/python:1-3.12-bullseye",
-	"features": {
-		"ghcr.io/devcontainers/features/docker-outside-of-docker:1": {},
-		"ghcr.io/devcontainers-extra/features/poetry:2": {},
-		"ghcr.io/devcontainers/features/node:1": {},
-	},
-	"postCreateCommand": ".devcontainer/setup.sh",
-	"runArgs": ["--add-host=host.docker.internal:host-gateway"],
-	"containerEnv": {
-		"DOCKER_HOST_ADDR": "host.docker.internal"
-	},
-}
--- a/.devcontainer/setup.sh
+++ b/.devcontainer/setup.sh
@@ -1,14 +0,0 @@
-#!/bin/bash
-
-# Mark the current repository as safe for Git to prevent "dubious ownership" errors,
-# which can occur in containerized environments when directory ownership doesn't match the current user.
-git config --global --add safe.directory "$(realpath .)"
-
-# Install `nc`
-sudo apt update && sudo apt install netcat -y
-
-# Install `uv` and `uvx`
-wget -qO- https://astral.sh/uv/install.sh | sh
-
-# Do common setup tasks
-source .openhands/setup.sh
--- a/.dockerignore
+++ b/.dockerignore
@@ -1,23 +1,5 @@
-# NodeJS
 frontend/node_modules
-
-# Configuration (except pyproject.toml)
-*.ini
-*.toml
-!pyproject.toml
-*.yml
-
-# Documentation (except README.md)
-*.md
-!README.md
-
-# Hidden files and directories
-.*
-__pycache__
-
-# Unneded files and directories
-/dev_config/
-/docs/
-/evaluation/
-/tests/
-CITATION.cff
+config.toml
+.envrc
+.env
+.git
--- a/.editorconfig
+++ b/.editorconfig
@@ -1,5 +0,0 @@
-[*]
-# force *nix line endings so files don't look modified in container run from Windows clone
-end_of_line = lf
-trim_trailing_whitespace = true
-insert_final_newline = true
--- a/.gitattributes
+++ b/.gitattributes
@@ -1,8 +1 @@
 *.ipynb linguist-vendored
-
-# force *nix line endings so files don't look modified in container run from Windows clone
-* text eol=lf
-# Git incorrectly thinks some media is text
-*.png -text
-*.gif -text
-*.mp4 -text
--- a/.github/.codecov.yml
+++ b/.github/.codecov.yml
@@ -0,0 +1,19 @@
+codecov:
+  notify:
+    wait_for_ci: true
+    # our project is large, so 6 builds are typically uploaded. this waits till 5/6
+    # See https://docs.codecov.com/docs/notifications#section-preventing-notifications-until-after-n-builds
+    after_n_builds: 5
+
+coverage:
+  status:
+    patch:
+      default:
+        threshold: 100% # allow patch coverage to be lower than project coverage by any amount
+    project:
+      default:
+        threshold: 5% # allow project coverage to drop at most 5%
+
+comment: false
+github_checks:
+    annotations: false
--- a/.github/ISSUE_TEMPLATE/bug_template.yml
+++ b/.github/ISSUE_TEMPLATE/bug_template.yml
@@ -1,166 +1,75 @@
 name: Bug
-description: Report a problem with OpenHands
+description: Report a problem with OpenDevin
 title: '[Bug]: '
 labels: ['bug']
 body:
  - type: markdown
    attributes:
-      value: |
-        ## Thank you for reporting a bug! 🐛
-
-        **Please fill out all required fields.** Issues missing critical information (version, installation method, reproduction steps, etc.) will be delayed or closed until complete details are provided.
-
-        Clear, detailed reports help us resolve issues faster.
+      value: Thank you for taking the time to fill out this bug report. We greatly appreciate your effort to complete this template fully. Please provide as much information as possible to help us understand and address the issue effectively.

  - type: checkboxes
    attributes:
      label: Is there an existing issue for the same bug?
-      description: Please search existing issues before creating a new one. If found, react or comment to the duplicate issue instead of making a new one.
+      description: Please check if an issue already exists for the bug you encountered.
      options:
-      - label: I have searched existing issues and this is not a duplicate.
+      - label: I have checked the troubleshooting document at https://opendevin.github.io/OpenDevin/modules/usage/troubleshooting
+        required: true
+      - label: I have checked the existing issues.
        required: true

  - type: textarea
    id: bug-description
    attributes:
-      label: Bug Description
-      description: Clearly describe what went wrong. Be specific and concise.
-      placeholder: Example - "When I run a Python task, OpenHands crashes after 30 seconds with a connection timeout error."
+      label: Describe the bug
+      description: Provide a short description of the problem.
    validations:
      required: true

  - type: textarea
-    id: expected-behavior
+    id: current-version
    attributes:
-      label: Expected Behavior
-      description: What did you expect to happen?
-      placeholder: Example - "OpenHands should execute the Python script and return results."
+      label: Current OpenDevin version
+      description: What version of OpenDevin are you using? If you're running in docker, tell us the tag you're using (e.g. ghcr.io/opendevin/opendevin:0.3.1).
+      render: bash
    validations:
-      required: false
+      required: true

  - type: textarea
-    id: actual-behavior
+    id: config
    attributes:
-      label: Actual Behavior
-      description: What actually happened?
-      placeholder: Example - "Connection timed out after 30 seconds, task failed with error code 500."
+      label: Installation and Configuration
+      description: Please provide any commands you ran and any configuration (redacting API keys)
+      render: bash
    validations:
-      required: false
+      required: true

  - type: textarea
-    id: reproduction-steps
+    id: model-agent
    attributes:
-      label: Steps to Reproduce
-      description: Provide clear, step-by-step instructions to reproduce the bug.
+      label: Model and Agent
+      description: What model and agent are you using? You can see these settings in the UI by clicking the settings wheel.
      placeholder: |
-        1. Install OpenHands using Docker
-        2. Configure with Claude 3.5 Sonnet
-        3. Run command: `openhands run "write a python script"`
-        4. Wait 30 seconds
-        5. Error appears
-    validations:
-      required: false
+        - Model:
+        - Agent:

-  - type: dropdown
-    id: installation
-    attributes:
-      label: OpenHands Installation Method
-      description: How are you running OpenHands?
-      options:
-        - CLI (uv tool install)
-        - CLI (executable binary)
-        - CLI (Docker)
-        - Local GUI (Docker web interface)
-        - OpenHands Cloud (app.all-hands.dev)
-        - SDK (Python library)
-        - Development workflow
-        - Other
-      default: 0
-    validations:
-      required: false
-
-  - type: input
-    id: installation-other
-    attributes:
-      label: If you selected "Other", please specify
-      description: Describe your installation method
-      placeholder: ex. Custom Kubernetes deployment, pip install from source, etc.
-
-  - type: input
-    id: openhands-version
-    attributes:
-      label: OpenHands Version
-      description: What version are you using? Find this in settings or by running `openhands --version`
-      placeholder: ex. 0.9.8, main, commit hash, etc.
-    validations:
-      required: false
-
-  - type: checkboxes
-    id: version-confirmation
-    attributes:
-      label: Version Confirmation
-      description: Bugs on older versions may already be fixed. Please upgrade before submitting.
-      options:
-        - label: "I have confirmed this bug exists on the LATEST version of OpenHands"
-          required: false
-
-  - type: input
-    id: model-name
-    attributes:
-      label: Model Name
-      description: Which LLM model are you using?
-      placeholder: ex. gpt-4o, claude-3-5-sonnet-20241022, openrouter/deepseek-r1, etc.
-    validations:
-      required: false
-
-  - type: dropdown
-    id: os
+  - type: textarea
+    id: os-version
    attributes:
      label: Operating System
-      options:
-        - MacOS
-        - Linux
-        - WSL on Windows
-        - Windows (Docker Desktop)
-        - Other
-    validations:
-      required: false
-
-  - type: input
-    id: browser
-    attributes:
-      label: Browser (if using web UI)
-      description: |
-        If applicable, which browser and version?
-
-      placeholder: ex. Chrome 131, Firefox 133, Safari 17.2
+      description: What Operating System are you using? Linux, Mac OS, WSL on Windows

  - type: textarea
-    id: logs
+    id: repro-steps
    attributes:
-      label: Logs and Error Messages
-      description: |
-        **Paste relevant logs, error messages, or stack traces.** Use code blocks (```) for formatting.
-
-        LLM logs are in `logs/llm/default/`. Include timestamps if errors occurred at a specific time.
+      label: Reproduction Steps
+      description: Please list the steps to reproduce the issue.
      placeholder: |
-        ```
-        Paste error logs here
-        ```
+        1.
+        2.
+        3.

  - type: textarea
    id: additional-context
    attributes:
-      label: Screenshots and Additional Context
-      description: |
-        Add screenshots, videos, runtime environment, or other context that helps explain the issue.
-
-        💡 **Share conversation history:** In the OpenHands chat UI, click the 👎 or 👍 button (above the message input) to generate a shareable link to your conversation.
-
-      placeholder: Drag and drop screenshots here, paste links, or add additional context.
-
-  - type: markdown
-    attributes:
-      value: |
-        ---
-        **Note:** Issues with incomplete information may be closed or deprioritized. Maintainers and community members have limited bandwidth and prioritize well-documented bugs that are easier to reproduce and fix. Thank you for your understanding!
+      label: Logs, Errors, Screenshots, and Additional Context
+      description: If you want to share the chat history you can click the thumbs-down (👎) button above the input field and you will get a shareable link (you can also click thumbs up when things are going well of course!). LLM logs will be stored in the `logs/llm/default` folder. Please add any additional context about the problem here.
--- a/.github/ISSUE_TEMPLATE/config.yml
+++ b/.github/ISSUE_TEMPLATE/config.yml
@@ -1,2 +0,0 @@
-# disable blank issue creation
-blank_issues_enabled: false
--- a/.github/ISSUE_TEMPLATE/feature_request.md
+++ b/.github/ISSUE_TEMPLATE/feature_request.md
@@ -0,0 +1,18 @@
+---
+name: Feature Request
+about: Suggest an idea for OpenDevin features
+title: ''
+labels: 'enhancement'
+assignees: ''
+
+---
+
+**What problem or use case are you trying to solve?**
+
+**Describe the UX of the solution you'd like**
+
+**Do you have thoughts on the technical implementation?**
+
+**Describe alternatives you've considered**
+
+**Additional context**
--- a/.github/ISSUE_TEMPLATE/feature_request.yml
+++ b/.github/ISSUE_TEMPLATE/feature_request.yml
@@ -1,105 +0,0 @@
-name: Feature Request or Enhancement
-description: Suggest a new feature or improvement for OpenHands
-title: '[Feature]: '
-labels: ['enhancement']
-body:
-  - type: markdown
-    attributes:
-      value: |
-        ## Thank you for suggesting a feature! 💡
-
-        **Please provide detailed information.** Vague or low-effort requests may be closed. Well-documented feature requests with strong community support are more likely to be added to the roadmap.
-
-  - type: checkboxes
-    attributes:
-      label: Is there an existing feature request for this?
-      description: Please search existing issues and feature requests before creating a new one. If found, react or comment to the duplicate issue instead of making a new one.
-      options:
-      - label: I have searched existing issues and feature requests, and this is not a duplicate.
-        required: true
-
-  - type: textarea
-    id: problem-statement
-    attributes:
-      label: Problem or Use Case
-      description: What problem are you trying to solve? What use case would this feature enable?
-      placeholder: |
-        Example - "As a developer working on large codebases, I need to search across multiple files simultaneously. Currently, I have to search file-by-file which is time-consuming and inefficient."
-    validations:
-      required: true
-
-  - type: textarea
-    id: proposed-solution
-    attributes:
-      label: Proposed Solution
-      description: Describe your ideal solution. What should this feature do? How should it work?
-      placeholder: |
-        Example - "Add a global search feature that allows searching across all files in the workspace. Results should show file name, line number, and context around matches. Include regex support and filtering options."
-    validations:
-      required: true
-
-  - type: textarea
-    id: alternatives
-    attributes:
-      label: Alternatives Considered
-      description: Have you considered any alternative solutions or workarounds? What are their limitations?
-      placeholder: Example - "I tried using grep in the terminal, but it's not integrated with the UI and doesn't provide click-to-navigate functionality."
-
-  - type: dropdown
-    id: priority
-    attributes:
-      label: Priority / Severity
-      description: How important is this feature to your workflow?
-      options:
-        - "Critical - Blocking my work, no workaround available"
-        - "High - Significant impact on productivity"
-        - "Medium - Would improve experience"
-        - "Low - Nice to have"
-      default: 2
-    validations:
-      required: true
-
-  - type: dropdown
-    id: scope
-    attributes:
-      label: Estimated Scope
-      description: To the best of your knowledge, how complex do you think this feature would be to implement?
-      options:
-        - "Small - UI tweak, config option, or minor change"
-        - "Medium - New feature with moderate complexity"
-        - "Large - Significant feature requiring architecture changes"
-        - "Unknown - Not sure about the technical complexity"
-      default: 3
-
-  - type: dropdown
-    id: feature-area
-    attributes:
-      label: Feature Area
-      description: Which part of OpenHands does this feature relate to? If you select "Other", please specify the area in the Additional Context section below.
-      options:
-        - "Agent / AI behavior"
-        - "User Interface / UX"
-        - "CLI / Command-line interface"
-        - "File system / Workspace management"
-        - "Configuration / Settings"
-        - "Integrations (GitHub, GitLab, etc.)"
-        - "Performance / Optimization"
-        - "Documentation"
-        - "Other"
-    validations:
-      required: true
-
-  - type: textarea
-    id: technical-details
-    attributes:
-      label: Technical Implementation Ideas (Optional)
-      description: If you have technical expertise, share implementation ideas, API suggestions, or relevant technical details.
-      placeholder: |
-        Example - "Could use ripgrep library for fast search. Expose results via /api/search endpoint. Frontend can use virtualized list for rendering large result sets."
-
-  - type: textarea
-    id: additional-context
-    attributes:
-      label: Additional Context
-      description: Add any other context, screenshots, mockups, or examples that help illustrate this feature request.
-      placeholder: Drag and drop screenshots, mockups, or links here.
--- a/.github/ISSUE_TEMPLATE/technical_proposal.md
+++ b/.github/ISSUE_TEMPLATE/technical_proposal.md
@@ -0,0 +1,18 @@
+---
+name: Technical Proposal
+about: Propose a new architecture or technology
+title: ''
+labels: 'proposal'
+assignees: ''
+
+---
+
+**Summary**
+
+**Motivation**
+
+**Technical Design**
+
+**Alternatives to Consider**
+
+**Additional context**
--- a/.github/actions/docker-image-tags/action.yml
+++ b/.github/actions/docker-image-tags/action.yml
@@ -1,51 +0,0 @@
-name: Compute Docker image tags
-description: Produce the canonical OpenHands Docker tag set (ref name, short SHA, full SHA — each in bare and `sha-` prefixed form) for a given image, with optional suffix and extra raw tags.
-
-inputs:
-  image:
-    description: Fully qualified image name (e.g. ghcr.io/owner/openhands).
-    required: true
-  ref-name:
-    description: Git ref name to emit as a tag (e.g. main, pr-123, saas-rel-1.2.3).
-    required: true
-  suffix:
-    description: Suffix appended to every tag (e.g. -amd64, -nikolaik-arm64). Leave empty for base (multi-arch manifest) tags.
-    required: false
-    default: ""
-  extra-tags:
-    description: Additional newline-separated metadata-action tag rules (e.g. extra `type=raw,value=...` lines).
-    required: false
-    default: ""
-
-outputs:
-  tags:
-    description: Newline-separated list of fully qualified image tags.
-    value: ${{ steps.meta.outputs.tags }}
-  labels:
-    description: Image labels emitted by docker/metadata-action.
-    value: ${{ steps.meta.outputs.labels }}
-  version:
-    description: Sanitized version string (ref-name with any suffix applied). Safe to use in docker tags.
-    value: ${{ steps.meta.outputs.version }}
-
-runs:
-  using: composite
-  steps:
-    - name: Compute tags
-      id: meta
-      uses: docker/metadata-action@v6
-      env:
-        # Use the PR head SHA (not the merge SHA) for sha-prefixed tags.
-        DOCKER_METADATA_PR_HEAD_SHA: "true"
-      with:
-        images: ${{ inputs.image }}
-        flavor: |
-          latest=false
-          suffix=${{ inputs.suffix }}
-        tags: |
-          type=raw,value=${{ inputs.ref-name }}
-          type=sha,prefix=sha-
-          type=sha,prefix=
-          type=sha,format=long,prefix=sha-
-          type=sha,format=long,prefix=
-          ${{ inputs.extra-tags }}
--- a/.github/actions/docker-merge-manifest/action.yml
+++ b/.github/actions/docker-merge-manifest/action.yml
@@ -1,43 +0,0 @@
-name: Merge multi-arch Docker manifest
-description: Build a multi-arch manifest from per-arch image tags pushed by an earlier build step.
-
-inputs:
-  base-tags:
-    description: Newline-separated list of base tags (without architecture suffix).
-    required: true
-  archs:
-    description: Space-separated list of architectures (e.g. "amd64 arm64").
-    required: true
-
-runs:
-  using: composite
-  steps:
-    - name: Login to GHCR
-      uses: docker/login-action@v4
-      with:
-        registry: ghcr.io
-        username: ${{ github.repository_owner }}
-        password: ${{ github.token }}
-
-    - name: Set up Docker Buildx
-      uses: docker/setup-buildx-action@v3
-
-    - name: Create multi-arch manifests
-      shell: bash
-      env:
-        BASE_TAGS: ${{ inputs.base-tags }}
-        ARCHS: ${{ inputs.archs }}
-      run: |
-        while IFS= read -r tag; do
-          [[ -z "$tag" ]] && continue
-          sources=""
-          for arch in $ARCHS; do
-            if ! docker buildx imagetools inspect "${tag}-${arch}" > /dev/null 2>&1; then
-              echo "::error::Missing image ${tag}-${arch}"
-              exit 1
-            fi
-            sources+=" ${tag}-${arch}"
-          done
-          echo "Creating manifest for $tag from:$sources"
-          docker buildx imagetools create -t "$tag" $sources
-        done <<< "$BASE_TAGS"
--- a/.github/dependabot.yml
+++ b/.github/dependabot.yml
@@ -1,82 +1,22 @@
+# To get started with Dependabot version updates, you'll need to specify which
+# package ecosystems to update and where the package manifests are located.
+# Please see the documentation for all configuration options:
+# https://docs.github.com/code-security/dependabot/dependabot-version-updates/configuration-options-for-the-dependabot.yml-file
+
 version: 2
 updates:
-  - package-ecosystem: "pip"
-    directory: "/"
+  - package-ecosystem: "pip" # See documentation for possible values
+    directory: "/" # Location of package manifests
    schedule:
      interval: "daily"
-    open-pull-requests-limit: 5
-    groups:
-      # put packages in their own group if they have a history of breaking the build or needing to be reverted
-      pre-commit:
-        patterns:
-          - "pre-commit"
-      browsergym:
-        patterns:
-          - "browsergym*"
-      mcp-packages:
-        patterns:
-          - "mcp"
-      security-all:
-        applies-to: "security-updates"
-        patterns:
-          - "*"
-      version-all:
-        applies-to: "version-updates"
-        patterns:
-          - "*"
-
-  - package-ecosystem: "npm"
-    directory: "/frontend"
+    open-pull-requests-limit: 20
+  - package-ecosystem: "npm" # See documentation for possible values
+    directory: "/frontend" # Location of package manifests
    schedule:
      interval: "daily"
-    open-pull-requests-limit: 5
-    groups:
-      docusaurus:
-        patterns:
-          - "*docusaurus*"
-      eslint:
-        patterns:
-          - "*eslint*"
-      security-all:
-        applies-to: "security-updates"
-        patterns:
-          - "*"
-      version-all:
-        applies-to: "version-updates"
-        patterns:
-          - "*"
-
-  - package-ecosystem: "npm"
-    directory: "/docs"
+    open-pull-requests-limit: 20
+  - package-ecosystem: "npm" # See documentation for possible values
+    directory: "/docs" # Location of package manifests
    schedule:
-      interval: "weekly"
-      day: "wednesday"
-    open-pull-requests-limit: 5
-    groups:
-      docusaurus:
-        patterns:
-          - "*docusaurus*"
-      eslint:
-        patterns:
-          - "*eslint*"
-      security-all:
-        applies-to: "security-updates"
-        patterns:
-          - "*"
-      version-all:
-        applies-to: "version-updates"
-        patterns:
-          - "*"
-
-  - package-ecosystem: "github-actions"
-    directory: "/"
-    schedule:
-      interval: "weekly"
-    open-pull-requests-limit: 5
-
-  - package-ecosystem: "docker"
-    directories:
-      - "containers/*"
-    schedule:
-      interval: "weekly"
-    open-pull-requests-limit: 5
+      interval: "daily"
+    open-pull-requests-limit: 20
--- a/.github/pull_request_template.md
+++ b/.github/pull_request_template.md
@@ -1,46 +1,5 @@
-<!-- Keep this PR as draft until it is ready for review. -->
+**What is the problem that this fixes or functionality that this introduces? Does it fix any open issues?**

-<!-- AI/LLM agents: be concise and specific. Do not check the box below. -->
+**Give a brief summary of what the PR does, explaining any non-trivial design decisions**

- [ ] A human has tested these changes.
-
---
-
-## Why
-
-<!-- Describe problem, motivation, etc.-->
-
-## Summary
-
-<!-- 1-3 bullets describing what changed. -->
-
-
-## Issue Number
-<!-- Required if there is a relevant issue to this PR. -->
-
-## How to Test
-
-<!--
-Required. Share the steps for the reviewer to be able to test your PR. e.g. You can test by running `npm install` then `npm build dev`.
-
-If you could not test this, say why.
-->
-
-## Video/Screenshots
-
-<!--
-Provide a video or screenshots of testing your PR. e.g. you added a new feature to the gui, show us the video of you testing it successfully.
-
-->
-
-## Type
-
- [ ] Bug fix
- [ ] Feature
- [ ] Refactor
- [ ] Breaking change
- [ ] Docs / chore
-
-## Notes
-
-<!-- Optional: migrations, config changes, rollout concerns, follow-ups, or anything reviewers should know. -->
+**Other references**
--- a/.github/scripts/find_prs_between_commits.py
+++ b/.github/scripts/find_prs_between_commits.py
@@ -1,330 +0,0 @@
-#!/usr/bin/env python3
-"""
-Find all PRs that went in between two commits in the OpenHands/OpenHands repository.
-Handles cherry-picks and different merge strategies.
-
-This script is designed to run from within the OpenHands repository under .github/scripts:
-    .github/scripts/find_prs_between_commits.py
-
-Usage: find_prs_between_commits <older_commit> <newer_commit> [--repo <path>]
-"""
-
-import json
-import os
-import re
-import subprocess
-import sys
-from collections import defaultdict
-from pathlib import Path
-from typing import Optional
-
-
-def find_openhands_repo() -> Optional[Path]:
-    """
-    Find the OpenHands repository.
-    Since this script is designed to live in .github/scripts/, it assumes
-    the repository root is two levels up from the script location.
-    Tries:
-    1. Repository root (../../ from script location)
-    2. Current directory
-    3. Environment variable OPENHANDS_REPO
-    """
-    # Check repository root (assuming script is in .github/scripts/)
-    script_dir = Path(__file__).parent.absolute()
-    repo_root = (
-        script_dir.parent.parent
-    )  # Go up two levels: scripts -> .github -> repo root
-    if (repo_root / '.git').exists():
-        return repo_root
-
-    # Check current directory
-    if (Path.cwd() / '.git').exists():
-        return Path.cwd()
-
-    # Check environment variable
-    if 'OPENHANDS_REPO' in os.environ:
-        repo_path = Path(os.environ['OPENHANDS_REPO'])
-        if (repo_path / '.git').exists():
-            return repo_path
-
-    return None
-
-
-def run_git_command(cmd: list[str], repo_path: Path) -> str:
-    """Run a git command in the repository directory and return its output."""
-    try:
-        result = subprocess.run(
-            cmd, capture_output=True, text=True, check=True, cwd=str(repo_path)
-        )
-        return result.stdout.strip()
-    except subprocess.CalledProcessError as e:
-        print(f'Error running git command: {" ".join(cmd)}', file=sys.stderr)
-        print(f'Error: {e.stderr}', file=sys.stderr)
-        sys.exit(1)
-
-
-def extract_pr_numbers_from_message(message: str) -> set[int]:
-    """Extract PR numbers from commit message in any common format."""
-    # Match #12345 anywhere, including in patterns like (#12345) or "Merge pull request #12345"
-    matches = re.findall(r'#(\d+)', message)
-    return set(int(m) for m in matches)
-
-
-def get_commit_info(commit_hash: str, repo_path: Path) -> tuple[str, str, str]:
-    """Get commit subject, body, and author from a commit hash."""
-    subject = run_git_command(
-        ['git', 'log', '-1', '--format=%s', commit_hash], repo_path
-    )
-    body = run_git_command(['git', 'log', '-1', '--format=%b', commit_hash], repo_path)
-    author = run_git_command(
-        ['git', 'log', '-1', '--format=%an <%ae>', commit_hash], repo_path
-    )
-    return subject, body, author
-
-
-def get_commits_between(
-    older_commit: str, newer_commit: str, repo_path: Path
-) -> list[str]:
-    """Get all commit hashes between two commits."""
-    commits_output = run_git_command(
-        ['git', 'rev-list', f'{older_commit}..{newer_commit}'], repo_path
-    )
-
-    if not commits_output:
-        return []
-
-    return commits_output.split('\n')
-
-
-def get_pr_info_from_github(pr_number: int, repo_path: Path) -> Optional[dict]:
-    """Get PR information from GitHub API if GITHUB_TOKEN is available."""
-    try:
-        # Set up environment with GitHub token
-        env = os.environ.copy()
-        if 'GITHUB_TOKEN' in env:
-            env['GH_TOKEN'] = env['GITHUB_TOKEN']
-
-        result = subprocess.run(
-            [
-                'gh',
-                'pr',
-                'view',
-                str(pr_number),
-                '--json',
-                'number,title,author,mergedAt,baseRefName,headRefName,url',
-            ],
-            capture_output=True,
-            text=True,
-            check=True,
-            env=env,
-            cwd=str(repo_path),
-        )
-        return json.loads(result.stdout)
-    except (subprocess.CalledProcessError, FileNotFoundError, json.JSONDecodeError):
-        return None
-
-
-def find_prs_between_commits(
-    older_commit: str, newer_commit: str, repo_path: Path
-) -> dict[int, dict]:
-    """
-    Find all PRs that went in between two commits.
-    Returns a dictionary mapping PR numbers to their information.
-    """
-    print(f'Repository: {repo_path}', file=sys.stderr)
-    print('Finding PRs between commits:', file=sys.stderr)
-    print(f'  Older: {older_commit}', file=sys.stderr)
-    print(f'  Newer: {newer_commit}', file=sys.stderr)
-    print(file=sys.stderr)
-
-    # Verify commits exist
-    try:
-        run_git_command(['git', 'rev-parse', '--verify', older_commit], repo_path)
-        run_git_command(['git', 'rev-parse', '--verify', newer_commit], repo_path)
-    except SystemExit:
-        print('Error: One or both commits not found in repository', file=sys.stderr)
-        sys.exit(1)
-
-    # Extract PRs from the older commit itself (to exclude from results)
-    # These PRs are already included at or before the older commit
-    older_subject, older_body, _ = get_commit_info(older_commit, repo_path)
-    older_message = f'{older_subject}\n{older_body}'
-    excluded_prs = extract_pr_numbers_from_message(older_message)
-
-    if excluded_prs:
-        print(
-            f'Excluding PRs already in older commit: {", ".join(f"#{pr}" for pr in sorted(excluded_prs))}',
-            file=sys.stderr,
-        )
-        print(file=sys.stderr)
-
-    # Get all commits between the two
-    commits = get_commits_between(older_commit, newer_commit, repo_path)
-    print(f'Found {len(commits)} commits to analyze', file=sys.stderr)
-    print(file=sys.stderr)
-
-    # Extract PR numbers from all commits
-    pr_info: dict[int, dict] = {}
-    commits_by_pr: dict[int, list[str]] = defaultdict(list)
-
-    for commit_hash in commits:
-        subject, body, author = get_commit_info(commit_hash, repo_path)
-        full_message = f'{subject}\n{body}'
-
-        pr_numbers = extract_pr_numbers_from_message(full_message)
-
-        for pr_num in pr_numbers:
-            # Skip PRs that are already in the older commit
-            if pr_num in excluded_prs:
-                continue
-
-            commits_by_pr[pr_num].append(commit_hash)
-
-            if pr_num not in pr_info:
-                pr_info[pr_num] = {
-                    'number': pr_num,
-                    'first_commit': commit_hash[:8],
-                    'first_commit_subject': subject,
-                    'commits': [],
-                    'github_info': None,
-                }
-
-            pr_info[pr_num]['commits'].append(
-                {'hash': commit_hash[:8], 'subject': subject, 'author': author}
-            )
-
-    # Try to get additional info from GitHub API
-    print('Fetching additional info from GitHub API...', file=sys.stderr)
-    for pr_num in pr_info.keys():
-        github_info = get_pr_info_from_github(pr_num, repo_path)
-        if github_info:
-            pr_info[pr_num]['github_info'] = github_info
-
-    print(file=sys.stderr)
-
-    return pr_info
-
-
-def print_results(pr_info: dict[int, dict]):
-    """Print the results in a readable format."""
-    sorted_prs = sorted(pr_info.items(), key=lambda x: x[0])
-
-    print(f'{"=" * 80}')
-    print(f'Found {len(sorted_prs)} PRs')
-    print(f'{"=" * 80}')
-    print()
-
-    for pr_num, info in sorted_prs:
-        print(f'PR #{pr_num}')
-
-        if info['github_info']:
-            gh = info['github_info']
-            print(f'  Title: {gh["title"]}')
-            print(f'  Author: {gh["author"]["login"]}')
-            print(f'  URL: {gh["url"]}')
-            if gh.get('mergedAt'):
-                print(f'  Merged: {gh["mergedAt"]}')
-            if gh.get('baseRefName'):
-                print(f'  Base: {gh["baseRefName"]} ← {gh["headRefName"]}')
-        else:
-            print(f'  Subject: {info["first_commit_subject"]}')
-
-        # Show if this PR has multiple commits (cherry-picked or multiple commits)
-        commit_count = len(info['commits'])
-        if commit_count > 1:
-            print(
-                f'  ⚠️  Found {commit_count} commits (possible cherry-pick or multi-commit PR):'
-            )
-            for commit in info['commits'][:3]:  # Show first 3
-                print(f'      {commit["hash"]}: {commit["subject"][:60]}')
-            if commit_count > 3:
-                print(f'      ... and {commit_count - 3} more')
-        else:
-            print(f'  Commit: {info["first_commit"]}')
-
-        print()
-
-
-def main():
-    if len(sys.argv) < 3:
-        print('Usage: find_prs_between_commits <older_commit> <newer_commit> [options]')
-        print()
-        print('Arguments:')
-        print('  <older_commit>  The older commit hash (or ref)')
-        print('  <newer_commit>  The newer commit hash (or ref)')
-        print()
-        print('Options:')
-        print('  --json          Output results in JSON format')
-        print('  --repo <path>   Path to OpenHands repository (default: auto-detect)')
-        print()
-        print('Example:')
-        print(
-            '  find_prs_between_commits c79e0cd3c7a2501a719c9296828d7a31e4030585 35bddb14f15124a3dc448a74651a6592911d99e9'
-        )
-        print()
-        print('Repository Detection:')
-        print('  The script will try to find the OpenHands repository in this order:')
-        print('  1. --repo argument')
-        print('  2. Repository root (../../ from script location)')
-        print('  3. Current directory')
-        print('  4. OPENHANDS_REPO environment variable')
-        print()
-        print('Environment variables:')
-        print(
-            '  GITHUB_TOKEN    Optional. If set, will fetch additional PR info from GitHub API'
-        )
-        print('  OPENHANDS_REPO  Optional. Path to OpenHands repository')
-        sys.exit(1)
-
-    older_commit = sys.argv[1]
-    newer_commit = sys.argv[2]
-    json_output = '--json' in sys.argv
-
-    # Check for --repo argument
-    repo_path = None
-    if '--repo' in sys.argv:
-        repo_idx = sys.argv.index('--repo')
-        if repo_idx + 1 < len(sys.argv):
-            repo_path = Path(sys.argv[repo_idx + 1])
-            if not (repo_path / '.git').exists():
-                print(f'Error: {repo_path} is not a git repository', file=sys.stderr)
-                sys.exit(1)
-
-    # Auto-detect repository if not specified
-    if repo_path is None:
-        repo_path = find_openhands_repo()
-        if repo_path is None:
-            print('Error: Could not find OpenHands repository', file=sys.stderr)
-            print('Please either:', file=sys.stderr)
-            print(
-                '  1. Place this script in .github/scripts/ within the OpenHands repository',
-                file=sys.stderr,
-            )
-            print('  2. Run from the OpenHands repository directory', file=sys.stderr)
-            print(
-                '  3. Use --repo <path> to specify the repository location',
-                file=sys.stderr,
-            )
-            print('  4. Set OPENHANDS_REPO environment variable', file=sys.stderr)
-            sys.exit(1)
-
-    # Find PRs
-    pr_info = find_prs_between_commits(older_commit, newer_commit, repo_path)
-
-    if json_output:
-        # Output as JSON
-        print(json.dumps(pr_info, indent=2))
-    else:
-        # Print results in human-readable format
-        print_results(pr_info)
-
-        # Also print a simple list for easy copying
-        print(f'{"=" * 80}')
-        print('PR Numbers (for easy copying):')
-        print(f'{"=" * 80}')
-        sorted_pr_nums = sorted(pr_info.keys())
-        print(', '.join(f'#{pr}' for pr in sorted_pr_nums))
-
-
-if __name__ == '__main__':
-    main()
--- a/.github/scripts/update_pr_description.sh
+++ b/.github/scripts/update_pr_description.sh
@@ -1,57 +0,0 @@
-#!/bin/bash
-
-set -euxo pipefail
-
-# This script updates the PR description with commands to run the PR locally
-# It adds both Docker and uvx commands
-
-# Get the branch name for the PR
-BRANCH_NAME=$(gh pr view "$PR_NUMBER" --json headRefName --jq .headRefName)
-
-# Define the Docker command
-DOCKER_RUN_COMMAND="docker run -it --rm \
-  -p 3000:3000 \
-  -v /var/run/docker.sock:/var/run/docker.sock \
-  --add-host host.docker.internal:host-gateway \
-  --name openhands-app-${SHORT_SHA} \
-  docker.openhands.dev/openhands/openhands:${SHORT_SHA}"
-
-# Get the current PR body
-PR_BODY=$(gh pr view "$PR_NUMBER" --json body --jq .body)
-
-# Prepare the new PR body with both commands
-if echo "$PR_BODY" | grep -q "To run this PR locally, use the following command:"; then
-  # For existing PR descriptions, use a more robust approach
-  # Split the PR body at the "To run this PR locally" section and replace everything after it
-  BEFORE_SECTION=$(echo "$PR_BODY" | sed '/To run this PR locally, use the following command:/,$d')
-  NEW_PR_BODY=$(cat <<EOF
-${BEFORE_SECTION}
-
-To run this PR locally, use the following command:
-
-GUI with Docker:
-\`\`\`
-${DOCKER_RUN_COMMAND}
-\`\`\`
-EOF
-)
-else
-  # For new PR descriptions: use heredoc safely without indentation
-  NEW_PR_BODY=$(cat <<EOF
-$PR_BODY
-
---
-
-To run this PR locally, use the following command:
-
-GUI with Docker:
-\`\`\`
-${DOCKER_RUN_COMMAND}
-\`\`\`
-EOF
-)
-fi
-
-# Update the PR description
-echo "Updating PR description with Docker and uvx commands"
-gh pr edit "$PR_NUMBER" --body "$NEW_PR_BODY"
--- a/.github/workflows/_build-image.yml
+++ b/.github/workflows/_build-image.yml
@@ -1,116 +0,0 @@
-# Reusable workflow: build a multi-arch Docker image and publish a merged manifest.
-# Called per image from .github/workflows/ghcr-build.yml.
-name: Build and push multi-arch image
-
-on:
-  workflow_call:
-    inputs:
-      image:
-        description: Fully-qualified image name (e.g. "ghcr.io/all-hands-ai/openhands").
-        required: true
-        type: string
-      context:
-        description: Docker build context.
-        required: false
-        type: string
-        default: "."
-      dockerfile:
-        description: Path to the Dockerfile.
-        required: true
-        type: string
-      extra-build-args:
-        description: Additional build-args (newline-separated). OPENHANDS_BUILD_VERSION is added automatically.
-        required: false
-        type: string
-        default: ""
-      provenance:
-        description: Value passed to docker/build-push-action provenance.
-        required: false
-        type: boolean
-        default: false
-      sbom:
-        description: Value passed to docker/build-push-action sbom.
-        required: false
-        type: boolean
-        default: false
-      buildx-driver-opts:
-        description: Extra buildx driver-opts (e.g. "network=host" for enterprise).
-        required: false
-        type: string
-        default: ""
-
-env:
-  RELEVANT_SHA: ${{ github.event.pull_request.head.sha || github.sha }}
-  RELEVANT_REF_NAME: ${{ github.event.pull_request.number && format('pr-{0}', github.event.pull_request.number) || github.ref_name }}
-
-jobs:
-  build:
-    name: Build ${{ inputs.image }} (${{ matrix.arch }})
-    runs-on: ${{ matrix.arch == 'arm64' && 'ubuntu-24.04-arm' || 'ubuntu-22.04' }}
-    permissions:
-      contents: read
-      packages: write
-    strategy:
-      matrix:
-        arch: [amd64, arm64]
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-        with:
-          ref: ${{ github.event.pull_request.head.sha }}
-      - name: Login to GHCR
-        uses: docker/login-action@v4
-        with:
-          registry: ghcr.io
-          username: ${{ github.repository_owner }}
-          password: ${{ secrets.GITHUB_TOKEN }}
-      - name: Set up Docker Buildx
-        uses: docker/setup-buildx-action@v3
-        with:
-          driver-opts: ${{ inputs.buildx-driver-opts }}
-      - name: Compute per-arch tags
-        id: meta
-        uses: ./.github/actions/docker-image-tags
-        with:
-          image: ${{ inputs.image }}
-          ref-name: ${{ env.RELEVANT_REF_NAME }}
-          suffix: -${{ matrix.arch }}
-      - name: Build and push
-        uses: docker/build-push-action@v7
-        with:
-          context: ${{ inputs.context }}
-          file: ${{ inputs.dockerfile }}
-          push: true
-          tags: ${{ steps.meta.outputs.tags }}
-          labels: ${{ steps.meta.outputs.labels }}
-          platforms: linux/${{ matrix.arch }}
-          build-args: |
-            OPENHANDS_BUILD_VERSION=${{ env.RELEVANT_REF_NAME }}
-            ${{ inputs.extra-build-args }}
-          cache-from: |
-            type=registry,ref=${{ inputs.image }}:buildcache-${{ steps.meta.outputs.version }}
-            type=registry,ref=${{ inputs.image }}:buildcache-main-${{ matrix.arch }}
-          cache-to: type=registry,ref=${{ inputs.image }}:buildcache-${{ steps.meta.outputs.version }},mode=max
-          provenance: ${{ inputs.provenance }}
-          sbom: ${{ inputs.sbom }}
-
-  merge:
-    name: Merge ${{ inputs.image }} manifest
-    runs-on: ubuntu-22.04
-    needs: build
-    permissions:
-      packages: write
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-      - name: Compute base tags
-        id: meta_base
-        uses: ./.github/actions/docker-image-tags
-        with:
-          image: ${{ inputs.image }}
-          ref-name: ${{ env.RELEVANT_REF_NAME }}
-      - name: Merge manifests
-        uses: ./.github/actions/docker-merge-manifest
-        with:
-          base-tags: ${{ steps.meta_base.outputs.tags }}
-          archs: "amd64 arm64"
--- a/.github/workflows/check-package-versions.yml
+++ b/.github/workflows/check-package-versions.yml
@@ -1,65 +0,0 @@
-name: Check Package Versions
-
-on:
-  push:
-    branches: [main]
-  pull_request:
-  workflow_dispatch:
-
-jobs:
-  check-package-versions:
-    runs-on: ubuntu-latest
-
-    steps:
-      - name: Checkout repository
-        uses: actions/checkout@v6
-
-      - name: Set up Python
-        uses: actions/setup-python@v6
-        with:
-          python-version: "3.12"
-
-      - name: Check for any 'rev' fields in pyproject.toml
-        run: |
-          python - <<'PY'
-          import sys, tomllib, pathlib
-
-          path = pathlib.Path("pyproject.toml")
-          if not path.exists():
-              print("❌ ERROR: pyproject.toml not found")
-              sys.exit(1)
-
-          try:
-              data = tomllib.loads(path.read_text(encoding="utf-8"))
-          except Exception as e:
-              print(f"❌ ERROR: Failed to parse pyproject.toml: {e}")
-              sys.exit(1)
-
-          poetry = data.get("tool", {}).get("poetry", {})
-          sections = {
-              "dependencies": poetry.get("dependencies", {}),
-          }
-
-          errors = []
-
-          print("🔍 Checking for any dependencies with 'rev' fields...\n")
-          for section_name, deps in sections.items():
-              if not isinstance(deps, dict):
-                  continue
-
-              for pkg_name, cfg in deps.items():
-                  if isinstance(cfg, dict) and "rev" in cfg:
-                      msg = f"  ✖ {pkg_name} in [{section_name}] uses rev='{cfg['rev']}' (NOT ALLOWED)"
-                      print(msg)
-                      errors.append(msg)
-                  else:
-                      print(f"  • {pkg_name}: OK")
-
-          if errors:
-              print("\n❌ FAILED: Found dependencies using 'rev' fields:\n" + "\n".join(errors))
-              print("\nPlease use versioned releases instead, e.g.:")
-              print('  my-package = "1.0.0"')
-              sys.exit(1)
-
-          print("\n✅ SUCCESS: No 'rev' fields found. All dependencies are using proper versioned releases.")
-          PY
--- a/.github/workflows/check-version-consistency.yml
+++ b/.github/workflows/check-version-consistency.yml
@@ -1,122 +0,0 @@
-name: Check Version Consistency
-
-on:
-  push:
-    branches: [main]
-  pull_request:
-  workflow_dispatch:
-
-jobs:
-  check-version-consistency:
-    runs-on: ubuntu-latest
-
-    steps:
-      - name: Checkout repository
-        uses: actions/checkout@v6
-
-      - name: Set up Python
-        uses: actions/setup-python@v6
-        with:
-          python-version: "3.12"
-
-      - name: Check version and Docker image tag consistency
-        run: |
-          python - <<'PY'
-          import json
-          import re
-          import sys
-          import tomllib
-
-          errors = []
-          warnings = []
-
-          # ── 1. Extract the canonical version from pyproject.toml ──────────
-          with open("pyproject.toml", "rb") as f:
-              pyproject = tomllib.load(f)
-          version = pyproject["tool"]["poetry"]["version"]
-          major_minor = ".".join(version.split(".")[:2])
-          print(f"📦 pyproject.toml version: {version} (major.minor: {major_minor})")
-
-          # ── 2. Check frontend/package.json ────────────────────────────────
-          with open("frontend/package.json") as f:
-              pkg = json.load(f)
-          if pkg["version"] != version:
-              errors.append(
-                  f"frontend/package.json version is '{pkg['version']}', expected '{version}'"
-              )
-          else:
-              print(f"  ✔ frontend/package.json: {pkg['version']}")
-
-          # ── 3. Check frontend/package-lock.json (2 places) ───────────────
-          with open("frontend/package-lock.json") as f:
-              lock = json.load(f)
-          for key, val in [
-              ("root.version", lock.get("version")),
-              ('packages[""].version', lock.get("packages", {}).get("", {}).get("version")),
-          ]:
-              if val != version:
-                  errors.append(
-                      f"frontend/package-lock.json {key} is '{val}', expected '{version}'"
-                  )
-              else:
-                  print(f"  ✔ frontend/package-lock.json {key}: {val}")
-
-          # ── 4. Check compose files use agent-server images ─────────────────
-          # Both compose files should use ghcr.io/.../agent-server (not runtime).
-          # Agent-server tags use SDK version (e.g. "1.12.0-python") or commit
-          # hashes (e.g. "31536c8-python") — both are acceptable.
-          repo_pattern = re.compile(r"AGENT_SERVER_IMAGE_REPOSITORY[^}]*:-([^}]+)")
-          tag_pattern = re.compile(r"AGENT_SERVER_IMAGE_TAG:-([^}]+)")
-
-          for filepath in ["docker-compose.yml", "containers/dev/compose.yml"]:
-              try:
-                  with open(filepath) as f:
-                      content = f.read()
-              except FileNotFoundError:
-                  warnings.append(f"{filepath}: file not found")
-                  continue
-
-              repos = repo_pattern.findall(content)
-              tags = tag_pattern.findall(content)
-
-              if not repos:
-                  warnings.append(f"{filepath}: no AGENT_SERVER_IMAGE_REPOSITORY default found")
-              else:
-                  repo = repos[0]
-                  if "agent-server" not in repo:
-                      errors.append(
-                          f"{filepath}: AGENT_SERVER_IMAGE_REPOSITORY defaults to '{repo}', "
-                          f"expected an agent-server image (not runtime)"
-                      )
-                  else:
-                      print(f"  ✔ {filepath} image repository: {repo}")
-
-              if not tags:
-                  warnings.append(f"{filepath}: no AGENT_SERVER_IMAGE_TAG default found")
-              else:
-                  tag = tags[0]
-                  if not tag:
-                      errors.append(f"{filepath}: AGENT_SERVER_IMAGE_TAG default is empty")
-                  else:
-                      print(f"  ✔ {filepath} image tag: {tag}")
-
-          # ── 5. Report ─────────────────────────────────────────────────────
-          print()
-          if warnings:
-              print("⚠ Warnings:")
-              for w in warnings:
-                  print(f"  {w}")
-              print()
-
-          if errors:
-              print("❌ FAILED: Version inconsistencies found:\n")
-              for e in errors:
-                  print(f"  ✖ {e}")
-              print(
-                  "\nAll version numbers and Docker image tags must be consistent."
-                  "\nSee .agents/skills/update-sdk/SKILL.md for the full checklist."
-              )
-              sys.exit(1)
-          else:
-              print("✅ All version numbers and Docker image tags are consistent.")
-          PY
--- a/.github/workflows/deploy-docs.yml
+++ b/.github/workflows/deploy-docs.yml
@@ -0,0 +1,59 @@
+name: Deploy Docs to GitHub Pages
+
+on:
+  push:
+    branches:
+      - main
+  pull_request:
+    branches:
+      - main
+
+jobs:
+  build:
+    name: Build Docusaurus
+    runs-on: ubuntu-latest
+    if: github.repository == 'OpenDevin/OpenDevin'
+    steps:
+      - uses: actions/checkout@v4
+        with:
+          fetch-depth: 0
+      - uses: actions/setup-node@v4
+        with:
+          node-version: 18
+          cache: npm
+          cache-dependency-path: docs/package-lock.json
+      - name: Set up Python
+        uses: actions/setup-python@v5
+        with:
+          python-version: "3.11"
+
+      - name: Generate Python Docs
+        run: rm -rf docs/modules/python && pip install pydoc-markdown && pydoc-markdown
+      - name: Install dependencies
+        run: cd docs && npm ci
+      - name: Build website
+        run: cd docs && npm run build
+
+      - name: Upload Build Artifact
+        if: github.ref == 'refs/heads/main'
+        uses: actions/upload-pages-artifact@v3
+        with:
+          path: docs/build
+
+  deploy:
+    name: Deploy to GitHub Pages
+    needs: build
+    if: github.ref == 'refs/heads/main' && github.repository == 'OpenDevin/OpenDevin'
+    # Grant GITHUB_TOKEN the permissions required to make a Pages deployment
+    permissions:
+      pages: write # to deploy to Pages
+      id-token: write # to verify the deployment originates from an appropriate source
+    # Deploy to the github-pages environment
+    environment:
+      name: github-pages
+      url: ${{ steps.deployment.outputs.page_url }}
+    runs-on: ubuntu-latest
+    steps:
+      - name: Deploy to GitHub Pages
+        id: deployment
+        uses: actions/deploy-pages@v4
--- a/.github/workflows/dummy-agent-test.yml
+++ b/.github/workflows/dummy-agent-test.yml
@@ -0,0 +1,42 @@
+name: Run E2E test with dummy agent
+
+concurrency:
+  group: ${{ github.workflow }}-${{ github.ref }}
+  cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
+
+on:
+  push:
+    branches:
+    - main
+  pull_request:
+
+env:
+  PERSIST_SANDBOX : "false"
+
+jobs:
+  test:
+    runs-on: ubuntu-latest
+    steps:
+      - uses: actions/checkout@v4
+      - name: Set up Python
+        uses: actions/setup-python@v5
+        with:
+          python-version: '3.11'
+      - name: Set up environment
+        run: |
+          curl -sSL https://install.python-poetry.org | python3 -
+          poetry install --without evaluation
+          poetry run playwright install --with-deps chromium
+          wget https://huggingface.co/BAAI/bge-small-en-v1.5/raw/main/1_Pooling/config.json -P /tmp/llama_index/models--BAAI--bge-small-en-v1.5/snapshots/5c38ec7c405ec4b44b94cc5a9bb96e735b38267a/1_Pooling/
+      - name: Run tests
+        run: |
+          set -e
+          poetry run python opendevin/core/main.py -t "do a flip" -d ./workspace/ -c DummyAgent
+      - name: Check exit code
+        run: |
+          if [ $? -ne 0 ]; then
+            echo "Test failed"
+            exit 1
+          else
+            echo "Test passed"
+          fi
--- a/.github/workflows/enterprise-check-migrations.yml
+++ b/.github/workflows/enterprise-check-migrations.yml
@@ -1,52 +0,0 @@
-name: Enterprise Check Migrations
-
-on:
-  pull_request:
-    paths:
-      - 'enterprise/migrations/**'
-
-jobs:
-  check-sync:
-    runs-on: ubuntu-latest
-    steps:
-      - name: Checkout PR branch
-        uses: actions/checkout@v6
-        with:
-          ref: ${{ github.event.pull_request.head.sha }}
-          fetch-depth: 0
-
-
-      - name: Fetch base branch
-        run: git fetch origin ${{ github.event.pull_request.base.ref }}
-
-      - name: Check if base branch is ancestor of PR
-        id: check_up_to_date
-        shell: bash
-        run: |
-          BASE="origin/${{ github.event.pull_request.base.ref }}"
-          HEAD="${{ github.event.pull_request.head.sha }}"
-          if git merge-base --is-ancestor "$BASE" "$HEAD"; then
-            echo "We're up to date with base $BASE"
-            exit 0
-          else
-            echo "NOT up to date with base $BASE"
-            exit 1
-          fi
-
-      - name: Find Comment
-        uses: peter-evans/find-comment@v4
-        id: find-comment
-        with:
-          issue-number: ${{ github.event.pull_request.number }}
-          comment-author: 'github-actions[bot]'
-          body-includes: |
-            ⚠️ This PR contains **migrations**
-
-      - name: Comment warning on PR
-        uses: peter-evans/create-or-update-comment@v5
-        with:
-          issue-number: ${{ github.event.pull_request.number }}
-          comment-id: ${{ steps.find-comment.outputs.comment-id }}
-          edit-mode: replace
-          body: |
-            ⚠️ This PR contains **migrations**. Please synchronize before merging to prevent conflicts.
--- a/.github/workflows/fe-e2e-tests.yml
+++ b/.github/workflows/fe-e2e-tests.yml
@@ -1,49 +0,0 @@
-# Workflow that runs frontend e2e tests with Playwright
-name: Run Frontend E2E Tests
-
-on:
-  push:
-    branches:
-      - main
-  pull_request:
-    paths:
-      - "frontend/**"
-      - ".github/workflows/fe-e2e-tests.yml"
-
-concurrency:
-  group: ${{ github.workflow }}-${{ (github.head_ref && github.ref) || github.run_id }}
-  cancel-in-progress: true
-
-jobs:
-  fe-e2e-test:
-    name: FE E2E Tests
-    runs-on: ubuntu-22.04
-    strategy:
-      matrix:
-        node-version: [22]
-      fail-fast: true
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-      - name: Set up Node.js
-        uses: actions/setup-node@v4
-        with:
-          node-version: ${{ matrix.node-version }}
-          cache: 'npm'
-          cache-dependency-path: frontend/package-lock.json
-      - name: Install dependencies
-        working-directory: ./frontend
-        run: npm ci
-      - name: Install Playwright browsers
-        working-directory: ./frontend
-        run: npx playwright install --with-deps chromium
-      - name: Run Playwright tests
-        working-directory: ./frontend
-        run: npx playwright test --project=chromium
-      - name: Upload Playwright report
-        uses: actions/upload-artifact@v7
-        if: always()
-        with:
-          name: playwright-report
-          path: frontend/playwright-report/
-          retention-days: 30
--- a/.github/workflows/fe-unit-tests.yml
+++ b/.github/workflows/fe-unit-tests.yml
@@ -1,46 +0,0 @@
-# Workflow that runs frontend unit tests
-name: Run Frontend Unit Tests
-
-# * Always run on "main"
-# * Run on PRs that have changes in the "frontend" folder or this workflow
-on:
-  push:
-    branches:
-      - main
-  pull_request:
-    paths:
-      - "frontend/**"
-      - ".github/workflows/fe-unit-tests.yml"
-
-# If triggered by a PR, it will be in the same group. However, each commit on main will be in its own unique group
-concurrency:
-  group: ${{ github.workflow }}-${{ (github.head_ref && github.ref) || github.run_id }}
-  cancel-in-progress: true
-
-jobs:
-  # Run frontend unit tests
-  fe-test:
-    name: FE Unit Tests
-    runs-on: ubuntu-22.04
-    strategy:
-      matrix:
-        node-version: [22]
-      fail-fast: true
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-      - name: Set up Node.js
-        uses: actions/setup-node@v4
-        with:
-          node-version: ${{ matrix.node-version }}
-          cache: 'npm'
-          cache-dependency-path: frontend/package-lock.json
-      - name: Install dependencies
-        working-directory: ./frontend
-        run: npm ci
-      - name: Run TypeScript compilation
-        working-directory: ./frontend
-        run: npm run build
-      - name: Run tests and collect coverage
-        working-directory: ./frontend
-        run: npm run test:coverage
--- a/.github/workflows/ghcr-build.yml
+++ b/.github/workflows/ghcr-build.yml
@@ -1,68 +0,0 @@
-# Workflow that builds and pushes the OpenHands app and enterprise Docker images to ghcr.io.
-# Per-image build logic lives in .github/workflows/_build-image.yml.
-name: Docker
-
-on:
-  push:
-    branches:
-      - main
-      - "saas-rel-*"
-      - "oss-rel-*"
-  pull_request:
-  workflow_dispatch:
-    inputs:
-      reason:
-        description: "Reason for manual trigger"
-        required: true
-        default: ""
-
-# PR events share a group so pushes supersede each other; each commit on a release branch gets its own group.
-concurrency:
-  group: ${{ github.workflow }}-${{ (github.head_ref && github.ref) || github.run_id }}
-  cancel-in-progress: true
-
-jobs:
-  build_app:
-    name: App
-    if: github.event.pull_request.head.repo.fork != true
-    uses: ./.github/workflows/_build-image.yml
-    with:
-      image: ghcr.io/openhands/openhands
-      dockerfile: containers/app/Dockerfile
-
-  build_enterprise:
-    name: Enterprise
-    if: github.event.pull_request.head.repo.fork != true
-    needs: build_app
-    uses: ./.github/workflows/_build-image.yml
-    with:
-      image: ghcr.io/openhands/enterprise-server
-      dockerfile: enterprise/Dockerfile
-      extra-build-args: OPENHANDS_VERSION=sha-${{ github.event.pull_request.head.sha || github.sha }}
-      provenance: true
-      sbom: true
-      buildx-driver-opts: network=host
-
-  update_pr_description:
-    name: Update PR Description
-    if: github.event_name == 'pull_request' && !github.event.pull_request.head.repo.fork && github.actor != 'dependabot[bot]'
-    needs: build_app
-    runs-on: ubuntu-22.04
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-
-      - name: Get short SHA
-        id: short_sha
-        run: echo "SHORT_SHA=$(echo ${{ github.event.pull_request.head.sha }} | cut -c1-7)" >> "$GITHUB_OUTPUT"
-
-      - name: Update PR Description
-        env:
-          GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
-          PR_NUMBER: ${{ github.event.pull_request.number }}
-          REPO: ${{ github.repository }}
-          SHORT_SHA: ${{ steps.short_sha.outputs.SHORT_SHA }}
-        shell: bash
-        run: |
-          echo "Updating PR description with Docker and uvx commands"
-          bash "${GITHUB_WORKSPACE}/.github/scripts/update_pr_description.sh"
--- a/.github/workflows/ghcr.yml
+++ b/.github/workflows/ghcr.yml
@@ -0,0 +1,263 @@
+name: Build Publish and Test Docker Image
+
+concurrency:
+  group: ${{ github.workflow }}-${{ github.ref }}
+  cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
+
+on:
+  push:
+    branches:
+      - main
+    tags:
+      - '*'
+  pull_request:
+  workflow_dispatch:
+    inputs:
+      reason:
+        description: 'Reason for manual trigger'
+        required: true
+        default: ''
+
+jobs:
+  ghcr_build:
+    runs-on: ubuntu-latest
+
+    outputs:
+      tags: ${{ steps.capture-tags.outputs.tags }}
+
+    permissions:
+      contents: read
+      packages: write
+
+    strategy:
+      matrix:
+        image: ["sandbox", "opendevin"]
+        platform: ["amd64", "arm64"]
+
+    steps:
+      - name: Checkout
+        uses: actions/checkout@v4
+
+      - name: Free Disk Space (Ubuntu)
+        uses: jlumbroso/free-disk-space@main
+        with:
+          # this might remove tools that are actually needed,
+          # if set to "true" but frees about 6 GB
+          tool-cache: true
+          # all of these default to true, but feel free to set to
+          # "false" if necessary for your workflow
+          android: true
+          dotnet: true
+          haskell: true
+          large-packages: true
+          docker-images: false
+          swap-storage: true
+
+      - name: Set up QEMU
+        uses: docker/setup-qemu-action@v3
+
+      - name: Set up Docker Buildx
+        id: buildx
+        uses: docker/setup-buildx-action@v3
+
+      - name: Build and export image
+        id: build
+        run: ./containers/build.sh ${{ matrix.image }} ${{ github.repository_owner }} ${{ matrix.platform }}
+
+      - name: Capture tags
+        id: capture-tags
+        run: |
+          tags=$(cat tags.txt)
+          echo "tags=$tags"
+          echo "tags=$tags" >> $GITHUB_OUTPUT
+
+      - name: Upload Docker image as artifact
+        uses: actions/upload-artifact@v4
+        with:
+          name: ${{ matrix.image }}-docker-image-${{ matrix.platform }}
+          path: /tmp/${{ matrix.image }}_image_${{ matrix.platform }}.tar
+
+  test-for-sandbox:
+    name: Test for Sandbox
+    runs-on: ubuntu-latest
+    needs: ghcr_build
+    env:
+      PERSIST_SANDBOX: "false"
+    steps:
+      - uses: actions/checkout@v4
+
+      - name: Install poetry via pipx
+        run: pipx install poetry
+
+      - name: Set up Python
+        uses: actions/setup-python@v5
+        with:
+          python-version: "3.11"
+          cache: "poetry"
+
+      - name: Install Python dependencies using Poetry
+        run: make install-python-dependencies
+
+      - name: Download sandbox Docker image
+        uses: actions/download-artifact@v4
+        with:
+          name: sandbox-docker-image-amd64
+          path: /tmp/
+
+      - name: Load sandbox image and run sandbox tests
+        run: |
+          # Load the Docker image and capture the output
+          output=$(docker load -i /tmp/sandbox_image_amd64.tar)
+
+          # Extract the first image name from the output
+          image_name=$(echo "$output" | grep -oP 'Loaded image: \K.*' | head -n 1)
+
+          # Print the full name of the image
+          echo "Loaded Docker image: $image_name"
+
+          SANDBOX_CONTAINER_IMAGE=$image_name TEST_IN_CI=true poetry run pytest --cov=agenthub --cov=opendevin --cov-report=xml -s ./tests/unit/test_sandbox.py
+
+      - name: Upload coverage to Codecov
+        uses: codecov/codecov-action@v4
+        env:
+          CODECOV_TOKEN: ${{ secrets.CODECOV_TOKEN }}
+
+  integration-tests-on-linux:
+    name: Integration Tests on Linux
+    runs-on: ubuntu-latest
+    needs: ghcr_build
+    env:
+      PERSIST_SANDBOX: "false"
+    strategy:
+      fail-fast: false
+      matrix:
+        python-version: ["3.11"]
+        sandbox: ["ssh", "local"]
+    steps:
+      - uses: actions/checkout@v4
+
+      - name: Install poetry via pipx
+        run: pipx install poetry
+
+      - name: Set up Python
+        uses: actions/setup-python@v5
+        with:
+          python-version: ${{ matrix.python-version }}
+          cache: 'poetry'
+
+      - name: Install Python dependencies using Poetry
+        run: make install-python-dependencies
+
+      - name: Download sandbox Docker image
+        uses: actions/download-artifact@v4
+        with:
+          name: sandbox-docker-image-amd64
+          path: /tmp/
+
+      - name: Load sandbox image and run integration tests
+        env:
+          SANDBOX_BOX_TYPE: ${{ matrix.sandbox }}
+        run: |
+          # Load the Docker image and capture the output
+          output=$(docker load -i /tmp/sandbox_image_amd64.tar)
+
+          # Extract the first image name from the output
+          image_name=$(echo "$output" | grep -oP 'Loaded image: \K.*' | head -n 1)
+
+          # Print the full name of the image
+          echo "Loaded Docker image: $image_name"
+
+          SANDBOX_CONTAINER_IMAGE=$image_name TEST_IN_CI=true TEST_ONLY=true ./tests/integration/regenerate.sh
+
+      - name: Upload coverage to Codecov
+        uses: codecov/codecov-action@v4
+        env:
+          CODECOV_TOKEN: ${{ secrets.CODECOV_TOKEN }}
+
+  ghcr_push:
+    runs-on: ubuntu-latest
+    # don't push if integration tests or sandbox tests fail
+    needs: [ghcr_build, integration-tests-on-linux, test-for-sandbox]
+    if: github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/')
+
+    env:
+      tags: ${{ needs.ghcr_build.outputs.tags }}
+
+    permissions:
+      contents: read
+      packages: write
+
+    strategy:
+      matrix:
+        image: ["sandbox", "opendevin"]
+        platform: ["amd64", "arm64"]
+
+    steps:
+      - name: Checkout code
+        uses: actions/checkout@v4
+
+      - name: Login to GHCR
+        uses: docker/login-action@v2
+        with:
+          registry: ghcr.io
+          username: ${{ github.repository_owner }}
+          password: ${{ secrets.GITHUB_TOKEN }}
+
+      - name: Download Docker images
+        uses: actions/download-artifact@v4
+        with:
+          name: ${{ matrix.image }}-docker-image-${{ matrix.platform }}
+          path: /tmp/${{ matrix.platform }}
+
+      - name: Load images and push to registry
+        run: |
+          mv /tmp/${{ matrix.platform }}/${{ matrix.image }}_image_${{ matrix.platform }}.tar .
+          loaded_image=$(docker load -i ${{ matrix.image }}_image_${{ matrix.platform }}.tar | grep "Loaded image:" | head -n 1 | awk '{print $3}')
+          echo "loaded image = $loaded_image"
+          tags=$(echo ${tags} | tr ' ' '\n')
+          image_name=$(echo "ghcr.io/${{ github.repository_owner }}/${{ matrix.image }}" | tr '[:upper:]' '[:lower:]')
+          echo "image name = $image_name"
+          for tag in $tags; do
+            echo "tag = $tag"
+            docker tag $loaded_image $image_name:${tag}_${{ matrix.platform }}
+            docker push $image_name:${tag}_${{ matrix.platform }}
+          done
+
+  create_manifest:
+    runs-on: ubuntu-latest
+    needs: [ghcr_build, ghcr_push]
+    if: github.ref == 'refs/heads/main' || startsWith(github.ref, 'refs/tags/')
+
+    env:
+      tags: ${{ needs.ghcr_build.outputs.tags }}
+
+    strategy:
+      matrix:
+        image: ["sandbox", "opendevin"]
+
+    permissions:
+      contents: read
+      packages: write
+
+    steps:
+      - name: Checkout code
+        uses: actions/checkout@v4
+
+      - name: Login to GHCR
+        uses: docker/login-action@v2
+        with:
+          registry: ghcr.io
+          username: ${{ github.repository_owner }}
+          password: ${{ secrets.GITHUB_TOKEN }}
+
+      - name: Create and push multi-platform manifest
+        run: |
+          image_name=$(echo "ghcr.io/${{ github.repository_owner }}/${{ matrix.image }}" | tr '[:upper:]' '[:lower:]')
+          echo "image name = $image_name"
+          tags=$(echo ${tags} | tr ' ' '\n')
+          for tag in $tags; do
+            echo 'tag = $tag'
+            docker buildx imagetools create --tag $image_name:$tag \
+              $image_name:${tag}_amd64 \
+              $image_name:${tag}_arm64
+          done
--- a/.github/workflows/lint-fix.yml
+++ b/.github/workflows/lint-fix.yml
@@ -1,98 +0,0 @@
-name: Lint Fix
-
-on:
-  pull_request:
-    types: [labeled]
-
-jobs:
-  # Frontend lint fixes
-  lint-fix-frontend:
-    if: github.event.label.name == 'lint-fix'
-    name: Fix frontend linting issues
-    runs-on: ubuntu-22.04
-    permissions:
-      contents: write
-      pull-requests: write
-    steps:
-      - uses: actions/checkout@v6
-        with:
-          ref: ${{ github.head_ref }}
-          repository: ${{ github.event.pull_request.head.repo.full_name }}
-          fetch-depth: 0
-          token: ${{ secrets.GITHUB_TOKEN }}
-
-      - name: Install Node.js 22
-        uses: actions/setup-node@v4
-        with:
-          node-version: 22
-          cache: 'npm'
-          cache-dependency-path: frontend/package-lock.json
-      - name: Install frontend dependencies
-        working-directory: ./frontend
-        run: npm ci
-      - name: Generate i18n and route types
-        run: |
-          cd frontend
-          npm run make-i18n
-          npx react-router typegen || true
-
-      - name: Fix frontend lint issues
-        run: |
-          cd frontend
-          npm run lint:fix
-
-      # Commit and push changes if any
-      - name: Check for changes
-        id: git-check
-        run: |
-          git diff --quiet || echo "changes=true" >> $GITHUB_OUTPUT
-      - name: Commit and push if there are changes
-        if: steps.git-check.outputs.changes == 'true'
-        run: |
-          git config --local user.email "openhands@all-hands.dev"
-          git config --local user.name "OpenHands Bot"
-          git add -A
-          git commit -m "🤖 Auto-fix frontend linting issues" --no-verify
-          git push
-
-  # Python lint fixes
-  lint-fix-python:
-    if: github.event.label.name == 'lint-fix'
-    name: Fix Python linting issues
-    runs-on: ubuntu-22.04
-    permissions:
-      contents: write
-      pull-requests: write
-    steps:
-      - uses: actions/checkout@v6
-        with:
-          ref: ${{ github.head_ref }}
-          repository: ${{ github.event.pull_request.head.repo.full_name }}
-          fetch-depth: 0
-          token: ${{ secrets.GITHUB_TOKEN }}
-
-      - name: Set up python
-        uses: actions/setup-python@v5
-        with:
-          python-version: 3.12
-          cache: "pip"
-      - name: Install pre-commit
-        run: pip install pre-commit==3.7.0
-      - name: Fix python lint issues
-        run: |
-          # Run all pre-commit hooks and continue even if they modify files (exit code 1)
-          pre-commit run --config ./dev_config/python/.pre-commit-config.yaml --all-files || true
-
-      # Commit and push changes if any
-      - name: Check for changes
-        id: git-check
-        run: |
-          git diff --quiet || echo "changes=true" >> $GITHUB_OUTPUT
-      - name: Commit and push if there are changes
-        if: steps.git-check.outputs.changes == 'true'
-        run: |
-          git config --local user.email "openhands@all-hands.dev"
-          git config --local user.name "OpenHands Bot"
-          git add -A
-          git commit -m "🤖 Auto-fix Python linting issues" --no-verify
-          git push
--- a/.github/workflows/lint.yml
+++ b/.github/workflows/lint.yml
@@ -1,75 +1,50 @@
-# Workflow that runs lint on the frontend and python code
 name: Lint

-# The jobs in this workflow are required, so they must run at all times
-# Always run on "main"
-# Always run on PRs
+concurrency:
+  group: ${{ github.workflow }}-${{ github.ref }}
+  cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
+
 on:
  push:
    branches:
-      - main
+    - main
  pull_request:

-# If triggered by a PR, it will be in the same group. However, each commit on main will be in its own unique group
-concurrency:
-  group: ${{ github.workflow }}-${{ (github.head_ref && github.ref) || github.run_id }}
-  cancel-in-progress: true
-
 jobs:
-  # Run lint on the frontend code
  lint-frontend:
    name: Lint frontend
-    runs-on: ubuntu-22.04
+    runs-on: ubuntu-latest
    steps:
-      - uses: actions/checkout@v6
-      - name: Install Node.js 22
+      - uses: actions/checkout@v4
+
+      - name: Install Node.js 20
        uses: actions/setup-node@v4
        with:
-          node-version: 22
-          cache: 'npm'
-          cache-dependency-path: frontend/package-lock.json
+          node-version: 20
+
      - name: Install dependencies
-        working-directory: ./frontend
-        run: npm ci
-      - name: Lint, TypeScript compilation, and translation checks
+        run: |
+          cd frontend
+          npm install --frozen-lockfile
+
+      - name: Lint
        run: |
          cd frontend
          npm run lint
-          npm run make-i18n && npx tsc
-          npm run check-translation-completeness

-  # Run lint on the python code
  lint-python:
    name: Lint python
-    runs-on: ubuntu-22.04
+    runs-on: ubuntu-latest
    steps:
-      - uses: actions/checkout@v6
+      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - name: Set up python
        uses: actions/setup-python@v5
        with:
-          python-version: 3.12
-          cache: "pip"
+          python-version: 3.11
+          cache: 'pip'
      - name: Install pre-commit
        run: pip install pre-commit==3.7.0
      - name: Run pre-commit hooks
-        run: pre-commit run --all-files --show-diff-on-failure --config ./dev_config/python/.pre-commit-config.yaml
-
-  lint-enterprise-python:
-    name: Lint enterprise python
-    runs-on: ubuntu-22.04
-    steps:
-      - uses: actions/checkout@v6
-        with:
-          fetch-depth: 0
-      - name: Set up python
-        uses: actions/setup-python@v5
-        with:
-          python-version: 3.12
-          cache: "pip"
-      - name: Install pre-commit
-        run: pip install pre-commit==4.2.0
-      - name: Run pre-commit hooks
-        working-directory: ./enterprise
-        run: pre-commit run --all-files --show-diff-on-failure --config ./dev_config/python/.pre-commit-config.yaml
+        run: pre-commit run --files opendevin/**/* agenthub/**/* evaluation/**/* tests/**/* --show-diff-on-failure --config ./dev_config/python/.pre-commit-config.yaml
--- a/.github/workflows/npm-publish-ui.yml
+++ b/.github/workflows/npm-publish-ui.yml
@@ -1,108 +0,0 @@
-name: Publish OpenHands UI Package
-
-# * Always run on "main"
-# * Run on PRs that have changes in the "openhands-ui" folder or this workflow
-on:
-  push:
-    branches:
-      - main
-    paths:
-      - "openhands-ui/**"
-      - ".github/workflows/npm-publish-ui.yml"
-
-# If triggered by a PR, it will be in the same group. However, each commit on main will be in its own unique group
-concurrency:
-  group: npm-publish-ui
-  cancel-in-progress: false
-
-jobs:
-  check-version:
-    name: Check if version has changed
-    runs-on: ubuntu-22.04
-    defaults:
-      run:
-        shell: bash
-    outputs:
-      should-publish: ${{ steps.version-check.outputs.should-publish }}
-      current-version: ${{ steps.version-check.outputs.current-version }}
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-        with:
-          fetch-depth: 2 # Need previous commit to compare
-
-      - name: Check if version changed
-        id: version-check
-        run: |
-          # Get current version from package.json
-          CURRENT_VERSION=$(jq -r .version openhands-ui/package.json)
-          echo "current-version=$CURRENT_VERSION" >> $GITHUB_OUTPUT
-
-          # Check if package.json version changed in this commit
-          if git diff HEAD~1 HEAD --name-only | grep -q "openhands-ui/package.json"; then
-            # Check if the version field specifically changed
-            if git diff HEAD~1 HEAD openhands-ui/package.json | grep -q '"version"'; then
-              echo "Version changed in package.json, will publish"
-              echo "should-publish=true" >> $GITHUB_OUTPUT
-            else
-              echo "package.json changed but version did not change, skipping publish"
-              echo "should-publish=false" >> $GITHUB_OUTPUT
-            fi
-          else
-            echo "package.json did not change, skipping publish"
-            echo "should-publish=false" >> $GITHUB_OUTPUT
-          fi
-
-  publish:
-    name: Publish to npm
-    runs-on: ubuntu-22.04
-    needs: check-version
-    if: needs.check-version.outputs.should-publish == 'true'
-    defaults:
-      run:
-        shell: bash
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-
-      - name: Setup Bun
-        uses: oven-sh/setup-bun@v2
-        with:
-          bun-version-file: "openhands-ui/.bun-version"
-
-      - name: Install dependencies
-        working-directory: ./openhands-ui
-        run: bun install --frozen-lockfile
-
-      - name: Build package
-        working-directory: ./openhands-ui
-        run: bun run build
-
-      - name: Check if package already exists on npm
-        id: npm-check
-        working-directory: ./openhands-ui
-        run: |
-          PACKAGE_NAME=$(jq -r .name package.json)
-          VERSION="${{ needs.check-version.outputs.current-version }}"
-
-          # Check if this version already exists on npm
-          if npm view "$PACKAGE_NAME@$VERSION" version 2>/dev/null; then
-            echo "Version $VERSION already exists on npm, skipping publish"
-            echo "already-exists=true" >> $GITHUB_OUTPUT
-          else
-            echo "Version $VERSION does not exist on npm, proceeding with publish"
-            echo "already-exists=false" >> $GITHUB_OUTPUT
-          fi
-
-      - name: Setup npm authentication
-        if: steps.npm-check.outputs.already-exists == 'false'
-        run: |
-          echo "//registry.npmjs.org/:_authToken=${{ secrets.NPM_TOKEN }}" > ~/.npmrc
-
-      - name: Publish to npm
-        if: steps.npm-check.outputs.already-exists == 'false'
-        working-directory: ./openhands-ui
-        run: |
-          # The prepublishOnly script will run automatically and build the package
-          npm publish
-          echo "✅ Successfully published @openhands/ui@${{ needs.check-version.outputs.current-version }} to npm"
--- a/.github/workflows/pr-artifacts.yml
+++ b/.github/workflows/pr-artifacts.yml
@@ -1,136 +0,0 @@
---
-name: PR Artifacts
-
-on:
-    workflow_dispatch: # Manual trigger for testing
-    pull_request:
-        types: [opened, synchronize, reopened]
-        branches: [main]
-    pull_request_review:
-        types: [submitted]
-
-jobs:
-  # Auto-remove .pr/ directory when a reviewer approves
-    cleanup-on-approval:
-        concurrency:
-            group: cleanup-pr-artifacts-${{ github.event.pull_request.number }}
-            cancel-in-progress: false
-        if: github.event_name == 'pull_request_review' && github.event.review.state == 'approved'
-        runs-on: ubuntu-latest
-        permissions:
-            contents: write
-            pull-requests: write
-        steps:
-            - name: Check if fork PR
-              id: check-fork
-              run: |
-                  if [ "${{ github.event.pull_request.head.repo.full_name }}" != "${{ github.event.pull_request.base.repo.full_name }}" ]; then
-                    echo "is_fork=true" >> $GITHUB_OUTPUT
-                    echo "::notice::Fork PR detected - skipping auto-cleanup (manual removal required)"
-                  else
-                    echo "is_fork=false" >> $GITHUB_OUTPUT
-                  fi
-
-            - uses: actions/checkout@v6
-              if: steps.check-fork.outputs.is_fork == 'false'
-              with:
-                  ref: ${{ github.event.pull_request.head.ref }}
-                  token: ${{ secrets.OPENHANDS_BOT_GITHUB_PAT_PUBLIC }}
-
-            - name: Remove .pr/ directory
-              id: remove
-              if: steps.check-fork.outputs.is_fork == 'false'
-              run: |
-                  if [ -d ".pr" ]; then
-                    git config user.name "allhands-bot"
-                    git config user.email "allhands-bot@users.noreply.github.com"
-                    git rm -rf .pr/
-                    git commit -m "chore: Remove PR-only artifacts [automated]"
-                    git push || {
-                      echo "::error::Failed to push cleanup commit. Check branch protection rules."
-                      exit 1
-                    }
-                    echo "removed=true" >> $GITHUB_OUTPUT
-                    echo "::notice::Removed .pr/ directory"
-                  else
-                    echo "removed=false" >> $GITHUB_OUTPUT
-                    echo "::notice::No .pr/ directory to remove"
-                  fi
-
-            - name: Update PR comment after cleanup
-              if: steps.check-fork.outputs.is_fork == 'false' && steps.remove.outputs.removed == 'true'
-              uses: actions/github-script@v9
-              with:
-                  script: |
-                      const marker = '<!-- pr-artifacts-notice -->';
-                      const body = `${marker}
-                      ✅ **PR Artifacts Cleaned Up**
-
-                      The \`.pr/\` directory has been automatically removed.
-                      `;
-
-                      const { data: comments } = await github.rest.issues.listComments({
-                        owner: context.repo.owner,
-                        repo: context.repo.repo,
-                        issue_number: context.issue.number,
-                      });
-
-                      const existing = comments.find(c => c.body.includes(marker));
-                      if (existing) {
-                        await github.rest.issues.updateComment({
-                          owner: context.repo.owner,
-                          repo: context.repo.repo,
-                          comment_id: existing.id,
-                          body: body,
-                        });
-                      }
-
-  # Warn if .pr/ directory exists (will be auto-removed on approval)
-    check-pr-artifacts:
-        if: github.event_name == 'pull_request'
-        runs-on: ubuntu-latest
-        permissions:
-            contents: read
-            pull-requests: write
-        steps:
-            - uses: actions/checkout@v6
-
-            - name: Check for .pr/ directory
-              id: check
-              run: |
-                  if [ -d ".pr" ]; then
-                    echo "exists=true" >> $GITHUB_OUTPUT
-                    echo "::warning::.pr/ directory exists and will be automatically removed when the PR is approved. For fork PRs, manual removal is required before merging."
-                  else
-                    echo "exists=false" >> $GITHUB_OUTPUT
-                  fi
-
-            - name: Post or update PR comment
-              if: steps.check.outputs.exists == 'true'
-              uses: actions/github-script@v9
-              with:
-                  script: |
-                      const marker = '<!-- pr-artifacts-notice -->';
-                      const body = `${marker}
-                      📁 **PR Artifacts Notice**
-
-                      This PR contains a \`.pr/\` directory with PR-specific documents. This directory will be **automatically removed** when the PR is approved.
-
-                      > For fork PRs: Manual removal is required before merging.
-                      `;
-
-                      const { data: comments } = await github.rest.issues.listComments({
-                        owner: context.repo.owner,
-                        repo: context.repo.repo,
-                        issue_number: context.issue.number,
-                      });
-
-                      const existing = comments.find(c => c.body.includes(marker));
-                      if (!existing) {
-                        await github.rest.issues.createComment({
-                          owner: context.repo.owner,
-                          repo: context.repo.repo,
-                          issue_number: context.issue.number,
-                          body: body,
-                        });
-                      }
--- a/.github/workflows/pr-review-by-openhands.yml
+++ b/.github/workflows/pr-review-by-openhands.yml
@@ -1,70 +0,0 @@
---
-name: PR Review by OpenHands
-
-on:
-    # Use pull_request for same-repo PRs so workflow changes can self-verify in PRs.
-    pull_request:
-        types: [opened, ready_for_review, labeled, review_requested]
-    # Use pull_request_target for fork PRs.
-    # The bot token used here is intentionally scoped to PR review operations,
-    # so the remaining blast radius is bounded even though PR content is untrusted.
-    pull_request_target:
-        types: [opened, ready_for_review, labeled, review_requested]
-
-permissions:
-    contents: read
-    pull-requests: write
-    issues: write
-
-jobs:
-    pr-review:
-        # Run on same-repo PRs via pull_request and on fork PRs via pull_request_target.
-        # Trigger when one of the following conditions is met:
-        #   1. A new non-draft PR is opened by a non-first-time contributor, OR
-        #   2. A draft PR is converted to ready for review by a non-first-time contributor, OR
-        #   3. The 'review-this' label is added, OR
-        #   4. openhands-agent or all-hands-bot is requested as a reviewer
-        # Note: FIRST_TIME_CONTRIBUTOR and NONE PRs require manual trigger via label/reviewer request.
-        # Trigger logic:
-        #   1. Route same-repo PRs through `pull_request` and fork PRs through `pull_request_target`
-        #   2. Auto-trigger on `opened` / `ready_for_review` for non-first-time contributors
-        #   3. Always allow manual triggers via `review-this` or reviewer request
-        # The author association check is duplicated intentionally for both
-        # auto-triggered actions (`opened` and `ready_for_review`).
-        if: |
-            (
-                (
-                    github.event_name == 'pull_request' &&
-                    github.event.pull_request.head.repo.full_name == github.repository
-                ) ||
-                (
-                    github.event_name == 'pull_request_target' &&
-                    github.event.pull_request.head.repo.full_name != github.repository
-                )
-            ) &&
-            (
-                (github.event.action == 'opened' && github.event.pull_request.draft == false && github.event.pull_request.author_association != 'FIRST_TIME_CONTRIBUTOR' && github.event.pull_request.author_association != 'NONE') ||
-                (github.event.action == 'ready_for_review' && github.event.pull_request.author_association != 'FIRST_TIME_CONTRIBUTOR' && github.event.pull_request.author_association != 'NONE') ||
-                (github.event.action == 'labeled' && github.event.label.name == 'review-this') ||
-                (
-                    github.event.action == 'review_requested' &&
-                    (
-                        github.event.requested_reviewer.login == 'openhands-agent' ||
-                        github.event.requested_reviewer.login == 'all-hands-bot'
-                    )
-                )
-            )
-        concurrency:
-            group: pr-review-${{ github.event.pull_request.number }}
-            cancel-in-progress: true
-        runs-on: ubuntu-24.04
-        steps:
-            - name: Run PR Review
-              uses: OpenHands/extensions/plugins/pr-review@main
-              with:
-                  llm-model: litellm_proxy/claude-sonnet-4-5-20250929
-                  llm-base-url: https://llm-proxy.app.all-hands.dev
-                  review-style: roasted
-                  llm-api-key: ${{ secrets.LLM_API_KEY }}
-                  github-token: ${{ secrets.OPENHANDS_BOT_GITHUB_PAT_PUBLIC }}
-                  lmnr-api-key: ${{ secrets.LMNR_SKILLS_API_KEY }}
--- a/.github/workflows/pr-review-evaluation.yml
+++ b/.github/workflows/pr-review-evaluation.yml
@@ -1,85 +0,0 @@
---
-name: PR Review Evaluation
-
-# This workflow evaluates how well PR review comments were addressed.
-# It runs when a PR is closed to assess review effectiveness.
-#
-# Security note: pull_request_target is safe here because:
-# 1. Only triggers on PR close (not on code changes)
-# 2. Does not checkout PR code - only downloads artifacts from trusted workflow runs
-# 3. Runs evaluation scripts from the extensions repo, not from the PR
-
-on:
-    pull_request_target:
-        types: [closed]
-
-permissions:
-    contents: read
-    pull-requests: read
-
-jobs:
-    evaluate:
-        runs-on: ubuntu-24.04
-        env:
-            PR_NUMBER: ${{ github.event.pull_request.number }}
-            REPO_NAME: ${{ github.repository }}
-            PR_MERGED: ${{ github.event.pull_request.merged }}
-
-        steps:
-            - name: Download review trace artifact
-              id: download-trace
-              uses: dawidd6/action-download-artifact@v15
-              continue-on-error: true
-              with:
-                  workflow: pr-review-by-openhands.yml
-                  name: pr-review-trace-${{ github.event.pull_request.number }}
-                  path: trace-info
-                  search_artifacts: true
-                  if_no_artifact_found: warn
-
-            - name: Check if trace file exists
-              id: check-trace
-              run: |
-                  if [ -f "trace-info/laminar_trace_info.json" ]; then
-                    echo "trace_exists=true" >> $GITHUB_OUTPUT
-                    echo "Found trace file for PR #$PR_NUMBER"
-                  else
-                    echo "trace_exists=false" >> $GITHUB_OUTPUT
-                    echo "No trace file found for PR #$PR_NUMBER - skipping evaluation"
-                  fi
-
-            # Always checkout main branch for security - cannot test script changes in PRs
-            - name: Checkout extensions repository
-              if: steps.check-trace.outputs.trace_exists == 'true'
-              uses: actions/checkout@v6
-              with:
-                  repository: OpenHands/extensions
-                  path: extensions
-
-            - name: Set up Python
-              if: steps.check-trace.outputs.trace_exists == 'true'
-              uses: actions/setup-python@v6
-              with:
-                  python-version: '3.12'
-
-            - name: Install dependencies
-              if: steps.check-trace.outputs.trace_exists == 'true'
-              run: pip install lmnr
-
-            - name: Run evaluation
-              if: steps.check-trace.outputs.trace_exists == 'true'
-              env:
-                  # Script expects LMNR_PROJECT_API_KEY; org secret is named LMNR_SKILLS_API_KEY
-                  LMNR_PROJECT_API_KEY: ${{ secrets.LMNR_SKILLS_API_KEY }}
-                  GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
-              run: |
-                  python extensions/plugins/pr-review/scripts/evaluate_review.py \
-                      --trace-file trace-info/laminar_trace_info.json
-
-            - name: Upload evaluation logs
-              uses: actions/upload-artifact@v7
-              if: always() && steps.check-trace.outputs.trace_exists == 'true'
-              with:
-                  name: pr-review-evaluation-${{ github.event.pull_request.number }}
-                  path: '*.log'
-                  retention-days: 30
--- a/.github/workflows/py-tests.yml
+++ b/.github/workflows/py-tests.yml
@@ -1,125 +0,0 @@
-# Workflow that runs python tests
-name: Run Python Tests
-
-# The jobs in this workflow are required, so they must run at all times
-# * Always run on "main"
-# * Always run on PRs
-on:
-  push:
-    branches:
-      - main
-  pull_request:
-
-# If triggered by a PR, it will be in the same group. However, each commit on main will be in its own unique group
-concurrency:
-  group: ${{ github.workflow }}-${{ (github.head_ref && github.ref) || github.run_id }}
-  cancel-in-progress: true
-
-jobs:
-  # Run python tests on Linux
-  test-on-linux:
-    name: Python Tests on Linux
-    runs-on: ubuntu-24.04
-    env:
-      INSTALL_DOCKER: "0" # Set to '0' to skip Docker installation
-    strategy:
-      matrix:
-        python-version: ["3.12"]
-    permissions:
-      # For coverage comment and python-coverage-comment-action branch
-      pull-requests: write
-      contents: write
-    steps:
-      - uses: actions/checkout@v6
-      - name: Set up Docker Buildx
-        id: buildx
-        uses: docker/setup-buildx-action@v3
-      - name: Install tmux
-        run: sudo apt-get update && sudo apt-get install -y tmux
-      - name: Setup Node.js
-        uses: actions/setup-node@v4
-        with:
-          node-version: "22.x"
-          cache: 'npm'
-          cache-dependency-path: frontend/package-lock.json
-      - name: Install poetry via pipx
-        run: pipx install poetry
-      - name: Set up Python
-        uses: actions/setup-python@v5
-        with:
-          python-version: ${{ matrix.python-version }}
-          cache: "poetry"
-      - name: Install Python dependencies using Poetry
-        run: |
-          poetry install --with dev,test,runtime
-          poetry run pip install pytest-xdist
-          poetry run pip install pytest-rerunfailures
-      - name: Build Environment
-        run: make build
-      - name: Run Unit Tests
-        run: PYTHONPATH=".:$PYTHONPATH" poetry run pytest --forked -n auto -s ./tests/unit --cov=openhands --cov-branch
-        env:
-          COVERAGE_FILE: ".coverage.${{ matrix.python_version }}"
-      - name: Store coverage file
-        uses: actions/upload-artifact@v7
-        with:
-          name: coverage-openhands
-          path: |
-            .coverage.${{ matrix.python_version }}
-            .coverage.runtime.${{ matrix.python_version }}
-          include-hidden-files: true
-
-  test-enterprise:
-    name: Enterprise Python Unit Tests
-    runs-on: ubuntu-24.04
-    strategy:
-      matrix:
-        python-version: ["3.12"]
-    steps:
-      - uses: actions/checkout@v6
-      - name: Install poetry via pipx
-        run: pipx install poetry
-      - name: Set up Python
-        uses: actions/setup-python@v5
-        with:
-          python-version: ${{ matrix.python-version }}
-          cache: "poetry"
-      - name: Install Python dependencies using Poetry
-        working-directory: ./enterprise
-        run: poetry install --with dev,test
-      - name: Run Unit Tests
-        # Use base working directory for coverage paths to line up.
-        run: PYTHONPATH=".:$PYTHONPATH" poetry run --project=enterprise pytest --forked -n auto -s -p no:ddtrace -p no:ddtrace.pytest_bdd -p no:ddtrace.pytest_benchmark ./enterprise/tests/unit --cov=enterprise --cov-branch
-        env:
-          COVERAGE_FILE: ".coverage.enterprise.${{ matrix.python_version }}"
-      - name: Store coverage file
-        uses: actions/upload-artifact@v7
-        with:
-          name: coverage-enterprise
-          path: ".coverage.enterprise.${{ matrix.python_version }}"
-          include-hidden-files: true
-
-  coverage-comment:
-    name: Coverage Comment
-    if: github.event_name == 'pull_request'
-    runs-on: ubuntu-latest
-    needs: [test-on-linux, test-enterprise]
-
-    permissions:
-      pull-requests: write
-      contents: write
-    steps:
-      - uses: actions/checkout@v6
-
-      - uses: actions/download-artifact@v8
-        id: download
-        with:
-          pattern: coverage-*
-          merge-multiple: true
-
-      - name: Coverage comment
-        id: coverage_comment
-        uses: py-cov-action/python-coverage-comment-action@v3
-        with:
-          GITHUB_TOKEN: ${{ github.token }}
-          MERGE_COVERAGE_FILES: true
--- a/.github/workflows/pypi-release.yml
+++ b/.github/workflows/pypi-release.yml
@@ -1,40 +0,0 @@
-# Publishes the OpenHands PyPi package
-name: Publish PyPi Package
-
-on:
-  workflow_dispatch:
-    inputs:
-      reason:
-        description: "What are you publishing?"
-        required: true
-        type: choice
-        options:
-          - app server
-        default: app server
-  push:
-    tags:
-      - "*"
-
-jobs:
-  release:
-    runs-on: ubuntu-22.04
-    # Run when manually dispatched for "app server" OR for tag pushes that don't contain '-cli' and don't start with 'cloud-'
-    if: |
-      (github.event_name == 'workflow_dispatch' && github.event.inputs.reason == 'app server')
-      || (github.event_name == 'push' && startsWith(github.ref, 'refs/tags/') && !contains(github.ref, '-cli') && !startsWith(github.ref, 'refs/tags/cloud-'))
-    steps:
-      - uses: actions/checkout@v6
-      - uses: actions/setup-python@v5
-        with:
-          python-version: 3.12
-      - name: Install Poetry
-        uses: snok/install-poetry@v1.4.1
-        with:
-          virtualenvs-in-project: true
-          virtualenvs-path: ~/.virtualenvs
-      - name: Install Poetry Dependencies
-        run: poetry install --no-interaction --no-root
-      - name: Build poetry project
-        run: ./build.sh
-      - name: publish
-        run: poetry publish -u __token__ -p ${{ secrets.PYPI_TOKEN }}
--- a/.github/workflows/review-pr.yml
+++ b/.github/workflows/review-pr.yml
@@ -0,0 +1,81 @@
+name: Use OpenDevin to Review Pull Request
+
+on:
+  pull_request:
+    types: [synchronize, labeled]
+
+permissions:
+  contents: write
+  pull-requests: write
+
+jobs:
+  dogfood:
+    if: contains(github.event.pull_request.labels.*.name, 'review-this')
+    runs-on: ubuntu-latest
+    container:
+      image: ghcr.io/opendevin/opendevin
+      volumes:
+        - /var/run/docker.sock:/var/run/docker.sock
+
+    steps:
+    - name: install git, github cli
+      run: |
+        apt-get install -y git gh
+        git config --global --add safe.directory $PWD
+
+    - name: Checkout Repository
+      uses: actions/checkout@v4
+      with:
+        ref: ${{ github.event.pull_request.base.ref }} # check out the target branch
+
+    - name: Download Diff
+      run: |
+        curl -O "${{ github.event.pull_request.diff_url }}" -L
+
+    - name: Write Task File
+      run: |
+        echo "Your coworker wants to apply a pull request to this project. Read and review ${{ github.event.pull_request.number }}.diff file. Create a review-${{ github.event.pull_request.number }}.txt and write your concise comments and suggestions there." > task.txt
+        echo "" >> task.txt
+        echo "Title" >> task.txt
+        echo "${{ github.event.pull_request.title }}" >> task.txt
+        echo "" >> task.txt
+        echo "Description" >> task.txt
+        echo "${{ github.event.pull_request.body }}" >> task.txt
+        echo "" >> task.txt
+        echo "Diff file is: ${{ github.event.pull_request.number }}.diff" >> task.txt
+
+    - name: Set up environment
+      run: |
+        curl -sSL https://install.python-poetry.org | python3 -
+        export PATH="/github/home/.local/bin:$PATH"
+        poetry install --without evaluation
+        poetry run playwright install --with-deps chromium
+
+    - name: Run OpenDevin
+      env:
+        LLM_API_KEY: ${{ secrets.OPENAI_API_KEY }}
+        OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
+        SANDBOX_BOX_TYPE: ssh
+      run: |
+        # Append path to launch poetry
+        export PATH="/github/home/.local/bin:$PATH"
+        # Append path to correctly import package, note: must set pwd at first
+        export PYTHONPATH=$(pwd):$PYTHONPATH
+        WORKSPACE_MOUNT_PATH=$GITHUB_WORKSPACE poetry run python ./opendevin/core/main.py -i 50 -f task.txt -d $GITHUB_WORKSPACE
+        rm task.txt
+
+    - name: Check if review file is non-empty
+      id: check_file
+      run: |
+        ls -la
+        if [[ -s review-${{ github.event.pull_request.number }}.txt ]]; then
+          echo "non_empty=true" >> $GITHUB_OUTPUT
+        fi
+      shell: bash
+
+    - name: Create PR review if file is non-empty
+      env:
+        GH_TOKEN: ${{ github.token }}
+      if: steps.check_file.outputs.non_empty == 'true'
+      run: |
+        gh pr review ${{ github.event.pull_request.number }} --comment --body-file "review-${{ github.event.pull_request.number }}.txt"
--- a/.github/workflows/run-unit-tests.yml
+++ b/.github/workflows/run-unit-tests.yml
@@ -0,0 +1,138 @@
+name: Run Unit Tests
+
+concurrency:
+  group: ${{ github.workflow }}-${{ github.ref }}
+  cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
+
+on:
+  push:
+    branches:
+      - main
+    paths-ignore:
+      - '**/*.md'
+      - 'frontend/**'
+      - 'docs/**'
+      - 'evaluation/**'
+  pull_request:
+
+env:
+  PERSIST_SANDBOX : "false"
+
+jobs:
+  fe-test:
+    runs-on: ubuntu-latest
+
+    strategy:
+      matrix:
+        node-version: [20]
+
+    steps:
+      - name: Checkout
+        uses: actions/checkout@v4
+
+      - name: Set up Node.js
+        uses: actions/setup-node@v4
+        with:
+          node-version: ${{ matrix.node-version }}
+
+      - name: Install dependencies
+        working-directory: ./frontend
+        run: npm ci
+
+      - name: Run tests and collect coverage
+        working-directory: ./frontend
+        run: npm run test:coverage
+
+      - name: Upload coverage to Codecov
+        uses: codecov/codecov-action@v4
+        env:
+          CODECOV_TOKEN: ${{ secrets.CODECOV_TOKEN }}
+
+  test-on-macos:
+    name: Test on macOS
+    runs-on: macos-12
+    env:
+      INSTALL_DOCKER: "1" # Set to '0' to skip Docker installation
+    strategy:
+      matrix:
+        python-version: ["3.11"]
+
+    steps:
+      - uses: actions/checkout@v4
+
+      - name: Install poetry via pipx
+        run: pipx install poetry
+
+      - name: Set up Python ${{ matrix.python-version }}
+        uses: actions/setup-python@v5
+        with:
+          python-version: ${{ matrix.python-version }}
+          cache: "poetry"
+
+      - name: Install Python dependencies using Poetry
+        run: poetry install
+
+      - name: Install & Start Docker
+        if: env.INSTALL_DOCKER == '1'
+        run: |
+          # Uninstall colima to upgrade to the latest version
+          if brew list colima &>/dev/null; then
+              brew uninstall colima
+              # unlinking colima dependency: go
+              brew uninstall go@1.21
+          fi
+          rm -rf ~/.colima ~/.lima
+          brew install --HEAD colima
+          brew services start colima
+          brew install docker
+          colima delete
+          colima start  --network-address --arch x86_64 --cpu=1 --memory=1
+
+          # For testcontainers to find the Colima socket
+          # https://github.com/abiosoft/colima/blob/main/docs/FAQ.md#cannot-connect-to-the-docker-daemon-at-unixvarrundockersock-is-the-docker-daemon-running
+          sudo ln -sf $HOME/.colima/default/docker.sock /var/run/docker.sock
+
+      - name: Build Environment
+        run: make build
+
+      - name: Run Tests
+        run: poetry run pytest --forked --cov=agenthub --cov=opendevin --cov-report=xml ./tests/unit -k "not test_sandbox"
+
+      - name: Upload coverage to Codecov
+        uses: codecov/codecov-action@v4
+        env:
+          CODECOV_TOKEN: ${{ secrets.CODECOV_TOKEN }}
+  test-on-linux:
+    name: Test on Linux
+    runs-on: ubuntu-latest
+    env:
+      INSTALL_DOCKER: "0" # Set to '0' to skip Docker installation
+    strategy:
+      matrix:
+        python-version: ["3.11"]
+
+    steps:
+      - uses: actions/checkout@v4
+
+      - name: Install poetry via pipx
+        run: pipx install poetry
+
+      - name: Set up Python
+        uses: actions/setup-python@v5
+        with:
+          python-version: ${{ matrix.python-version }}
+          cache: "poetry"
+
+      - name: Install Python dependencies using Poetry
+        run: poetry install --without evaluation
+
+      - name: Build Environment
+        run: make build
+
+      - name: Run Tests
+        run: poetry run pytest --forked --cov=agenthub --cov=opendevin --cov-report=xml ./tests/unit -k "not test_sandbox"
+
+      - name: Upload coverage to Codecov
+        uses: codecov/codecov-action@v4
+        env:
+          CODECOV_TOKEN: ${{ secrets.CODECOV_TOKEN }}
--- a/.github/workflows/solve-issue.yml
+++ b/.github/workflows/solve-issue.yml
@@ -0,0 +1,122 @@
+name: Use OpenDevin to Resolve GitHub Issue
+
+on:
+  issues:
+    types: [labeled]
+
+permissions:
+  contents: write
+  pull-requests: write
+  issues: write
+
+jobs:
+  dogfood:
+    if: github.event.label.name == 'solve-this'
+    runs-on: ubuntu-latest
+    container:
+      image: ghcr.io/opendevin/opendevin
+      volumes:
+        - /var/run/docker.sock:/var/run/docker.sock
+
+    steps:
+    - name: install git, github cli
+      run: apt-get install -y git gh
+
+    - name: Checkout Repository
+      uses: actions/checkout@v4
+
+    - name: Write Task File
+      env:
+        ISSUE_TITLE: ${{ github.event.issue.title }}
+        ISSUE_BODY: ${{ github.event.issue.body }}
+      run: |
+        echo "TITLE:" > task.txt
+        echo "${ISSUE_TITLE}" >> task.txt
+        echo "" >> task.txt
+        echo "BODY:" >> task.txt
+        echo "${ISSUE_BODY}" >> task.txt
+
+    - name: Set up environment
+      run: |
+        curl -sSL https://install.python-poetry.org | python3 -
+        export PATH="/github/home/.local/bin:$PATH"
+        poetry install --without evaluation
+        poetry run playwright install --with-deps chromium
+
+
+    - name: Run OpenDevin
+      env:
+        ISSUE_TITLE: ${{ github.event.issue.title }}
+        ISSUE_BODY: ${{ github.event.issue.body }}
+        LLM_API_KEY: ${{ secrets.OPENAI_API_KEY }}
+        OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
+        SANDBOX_BOX_TYPE: ssh
+      run: |
+        # Append path to launch poetry
+        export PATH="/github/home/.local/bin:$PATH"
+        # Append path to correctly import package, note: must set pwd at first
+        export PYTHONPATH=$(pwd):$PYTHONPATH
+        WORKSPACE_MOUNT_PATH=$GITHUB_WORKSPACE poetry run python ./opendevin/core/main.py -i 50 -f task.txt -d $GITHUB_WORKSPACE
+        rm task.txt
+
+    - name: Setup Git, Create Branch, and Commit Changes
+      run: |
+        # Setup Git configuration
+        git config --global --add safe.directory $PWD
+        git config --global user.name 'OpenDevin'
+        git config --global user.email 'OpenDevin@users.noreply.github.com'
+
+        # Create a unique branch name with a timestamp
+        BRANCH_NAME="fix/${{ github.event.issue.number }}-$(date +%Y%m%d%H%M%S)"
+
+        # Checkout new branch
+        git checkout -b $BRANCH_NAME
+
+        # Add all changes to staging, except task.txt
+        git add --all -- ':!task.txt'
+
+        # Commit the changes, if any
+        git commit -m "OpenDevin: Resolve Issue #${{ github.event.issue.number }}"
+        if [ $? -ne 0 ]; then
+          echo "No changes to commit."
+          exit 0
+        fi
+
+        # Push changes
+        git push --set-upstream origin $BRANCH_NAME
+
+    - name: Fetch Default Branch
+      env:
+        GH_TOKEN: ${{ github.token }}
+      run: |
+        # Fetch the default branch using gh cli
+        DEFAULT_BRANCH=$(gh repo view --json defaultBranchRef --jq .defaultBranchRef.name)
+        echo "Default branch is $DEFAULT_BRANCH"
+        echo "DEFAULT_BRANCH=$DEFAULT_BRANCH" >> $GITHUB_ENV
+
+    - name: Generate PR
+      env:
+        GH_TOKEN: ${{ github.token }}
+      run: |
+        # Create PR and capture URL
+        PR_URL=$(gh pr create \
+          --title "OpenDevin: Resolve Issue #2" \
+          --body "This PR was generated by OpenDevin to resolve issue #2" \
+          --repo "foragerr/OpenDevin" \
+          --head "${{ github.head_ref }}" \
+          --base "${{ env.DEFAULT_BRANCH }}" \
+          | grep -o 'https://github.com/[^ ]*')
+
+        # Extract PR number from URL
+        PR_NUMBER=$(echo "$PR_URL" | grep -o '[0-9]\+$')
+
+        # Set environment vars
+        echo "PR_URL=$PR_URL" >> $GITHUB_ENV
+        echo "PR_NUMBER=$PR_NUMBER" >> $GITHUB_ENV
+
+    - name: Post Comment
+      env:
+        GH_TOKEN: ${{ github.token }}
+      run: |
+        gh issue comment ${{ github.event.issue.number }} \
+          -b "OpenDevin raised [PR #${{ env.PR_NUMBER }}](${{ env.PR_URL }}) to resolve this issue."
--- a/.github/workflows/stale.yml
+++ b/.github/workflows/stale.yml
@@ -1,23 +1,29 @@
-# Workflow that marks issues and PRs with no activity for 30 days with "Stale" and closes them after 7 more days of no activity
 name: 'Close stale issues'
-
-# Runs every day at 01:30
 on:
  schedule:
    - cron: '30 1 * * *'

 jobs:
  stale:
-    runs-on: ubuntu-22.04
-    if: github.repository == 'OpenHands/OpenHands'
+    runs-on: ubuntu-latest
    steps:
-      - uses: actions/stale@v10
+      - uses: actions/stale@v9
        with:
-          stale-issue-message: 'This issue is stale because it has been open for 40 days with no activity. Remove the stale label or leave a comment, otherwise it will be closed in 10 days.'
-          stale-pr-message: 'This PR is stale because it has been open for 40 days with no activity. Remove the stale label or leave a comment, otherwise it will be closed in 10 days.'
-          days-before-stale: 40
-          exempt-issue-labels: roadmap,backlog,app-team
-          close-issue-message: 'This issue was automatically closed due to 50 days of inactivity. We do this to help keep the issues somewhat manageable and focus on active issues.'
-          close-pr-message: 'This PR was closed because it had no activity for 50 days. If you feel this was closed in error, and you would like to continue the PR, please resubmit or let us know.'
-          days-before-close: 10
-          operations-per-run: 300
+          # Aggressively close issues that have been explicitly labeled `age-out`
+          any-of-labels: age-out
+          stale-issue-message: 'This issue is stale because it has been open for 7 days with no activity. Remove stale label or comment or this will be closed in 1 day.'
+          close-issue-message: 'This issue was closed because it has been stalled for over 7 days with no activity.'
+          stale-pr-message: 'This PR is stale because it has been open for 7 days with no activity. Remove stale label or comment or this will be closed in 1 days.'
+          close-pr-message: 'This PR was closed because it has been stalled for over 7 days with no activity.'
+          days-before-stale: 7
+          days-before-close: 1
+
+      - uses: actions/stale@v9
+        with:
+          # Be more lenient with other issues
+          stale-issue-message: 'This issue is stale because it has been open for 30 days with no activity. Remove stale label or comment or this will be closed in 7 days.'
+          close-issue-message: 'This issue was closed because it has been stalled for over 30 days with no activity.'
+          stale-pr-message: 'This PR is stale because it has been open for 30 days with no activity. Remove stale label or comment or this will be closed in 7 days.'
+          close-pr-message: 'This PR was closed because it has been stalled for over 30 days with no activity.'
+          days-before-stale: 30
+          days-before-close: 7
--- a/.github/workflows/tag-image.yml
+++ b/.github/workflows/tag-image.yml
@@ -1,59 +0,0 @@
-# Adds a git-tag name to existing Docker images.
-# Triggered when a tag is pushed: finds the images built at the tag's commit
-# (tagged `sha-<full>`) and adds the tag name as an alias for the same manifest.
-# Semver tags (X.Y.Z) also get X.Y, X, and latest aliases.
-# No rebuild — pure registry-side retag via `docker buildx imagetools create`.
-name: Tag Docker images
-
-on:
-  push:
-    tags:
-      - "*"
-
-jobs:
-  retag:
-    runs-on: ubuntu-22.04
-    permissions:
-      packages: write
-    strategy:
-      matrix:
-        image:
-          - ghcr.io/openhands/openhands
-          - ghcr.io/openhands/enterprise-server
-    steps:
-      - name: Login to GHCR
-        uses: docker/login-action@v4
-        with:
-          registry: ghcr.io
-          username: ${{ github.repository_owner }}
-          password: ${{ secrets.GITHUB_TOKEN }}
-      - name: Set up Docker Buildx
-        uses: docker/setup-buildx-action@v3
-      - name: Compute tags
-        id: meta
-        uses: docker/metadata-action@v6
-        with:
-          images: ${{ matrix.image }}
-          flavor: latest=auto
-          tags: |
-            type=ref,event=tag
-            type=semver,pattern={{version}}
-            type=semver,pattern={{major}}.{{minor}}
-            type=semver,pattern={{major}}
-      - name: Add tags to existing image
-        env:
-          SRC: ${{ matrix.image }}:sha-${{ github.sha }}
-          TAGS: ${{ steps.meta.outputs.tags }}
-        shell: bash
-        run: |
-          set -euo pipefail
-          if ! docker buildx imagetools inspect "$SRC" > /dev/null 2>&1; then
-            echo "::error::Source image $SRC does not exist. The Docker workflow for commit ${{ github.sha }} may not have completed successfully. Re-run this workflow once the build finishes."
-            exit 1
-          fi
-          args=()
-          while IFS= read -r tag; do
-            [[ -z "$tag" ]] && continue
-            args+=(-t "$tag")
-          done <<< "$TAGS"
-          docker buildx imagetools create "${args[@]}" "$SRC"
--- a/.github/workflows/ui-build.yml
+++ b/.github/workflows/ui-build.yml
@@ -1,34 +0,0 @@
-name: Run UI Component Build
-
-# * Always run on "main"
-# * Run on PRs that have changes in the "openhands-ui" folder or this workflow
-on:
-  push:
-    branches:
-      - main
-  pull_request:
-    paths:
-      - 'openhands-ui/**'
-      -  '.github/workflows/ui-build.yml'
-
-# If triggered by a PR, it will be in the same group. However, each commit on main will be in its own unique group
-concurrency:
-  group: ${{ github.workflow }}-${{ (github.head_ref && github.ref) || github.run_id }}
-  cancel-in-progress: true
-
-jobs:
-  ui-build:
-    name: Build openhands-ui
-    runs-on: ubuntu-22.04
-    steps:
-      - name: Checkout
-        uses: actions/checkout@v6
-      - uses: oven-sh/setup-bun@v2
-        with:
-          bun-version-file: "openhands-ui/.bun-version"
-      - name: Install dependencies
-        working-directory: ./openhands-ui
-        run: bun install --frozen-lockfile
-      - name: Build package
-        working-directory:  ./openhands-ui
-        run: bun run build
--- a/.github/workflows/update-pyproject-version.yml
+++ b/.github/workflows/update-pyproject-version.yml
@@ -0,0 +1,48 @@
+name: Update pyproject.toml Version and Tags
+
+on:
+  release:
+    types:
+      - published
+
+jobs:
+  update-pyproject-and-tags:
+    runs-on: ubuntu-latest
+
+    steps:
+      - name: Checkout code
+        uses: actions/checkout@v4
+        with:
+          fetch-depth: 0  # Fetch all history for all branches and tags
+
+      - name: Set up Python
+        uses: actions/setup-python@v5
+        with:
+          python-version: "3.11"
+
+      - name: Install dependencies
+        run: |
+          python -m pip install --upgrade pip
+          pip install toml
+
+      - name: Get release tag
+        id: get_release_tag
+        run: echo "RELEASE_TAG=${GITHUB_REF#refs/tags/}" >> $GITHUB_ENV
+
+      - name: Update pyproject.toml with release tag
+        run: |
+          python -c "
+          import toml
+          with open('pyproject.toml', 'r') as f:
+              data = toml.load(f)
+          data['tool']['poetry']['version'] = '${{ env.RELEASE_TAG }}'
+          with open('pyproject.toml', 'w') as f:
+              toml.dump(data, f)
+          "
+
+      - name: Commit and push pyproject.toml changes
+        uses: stefanzweifel/git-auto-commit-action@v4
+        with:
+          commit_message: "Update pyproject.toml version to ${{ env.RELEASE_TAG }}"
+          branch: main
+          file_pattern: pyproject.toml
--- a/.github/workflows/welcome-good-first-issue.yml
+++ b/.github/workflows/welcome-good-first-issue.yml
@@ -1,51 +0,0 @@
-name: Welcome Good First Issue
-
-on:
-  issues:
-    types: [labeled]
-
-permissions:
-  issues: write
-
-jobs:
-  comment-on-good-first-issue:
-    if: github.event.label.name == 'good first issue'
-    runs-on: ubuntu-latest
-    steps:
-      - name: Check if welcome comment already exists
-        id: check_comment
-        uses: actions/github-script@v9
-        with:
-          result-encoding: string
-          script: |
-            const issueNumber = context.issue.number;
-            const comments = await github.rest.issues.listComments({
-              ...context.repo,
-              issue_number: issueNumber
-            });
-
-            const alreadyCommented = comments.data.some(
-              (comment) =>
-                comment.body.includes('<!-- auto-comment:good-first-issue -->')
-            );
-
-            return alreadyCommented ? 'true' : 'false';
-
-      - name: Leave welcome comment
-        if: steps.check_comment.outputs.result == 'false'
-        uses: actions/github-script@v9
-        with:
-          script: |
-            const repoUrl = `https://github.com/${context.repo.owner}/${context.repo.repo}`;
-
-            await github.rest.issues.createComment({
-              ...context.repo,
-              issue_number: context.issue.number,
-              body: "🙌 **Hey there, future contributor!** 🙌\n\n" +
-                    "This issue has been labeled as **good first issue**, which means it's a great place to get started with the OpenHands project.\n\n" +
-                    "If you're interested in working on it, feel free to! No need to ask for permission.\n\n" +
-                    "Be sure to check out our [development setup guide](" + repoUrl + "/blob/main/Development.md) to get your environment set up, and follow our [contribution guidelines](" + repoUrl + "/blob/main/CONTRIBUTING.md) when you're ready to submit a fix.\n\n" +
-                    "Feel free to join our developer community on [Slack](https://openhands.dev/joinslack). You can ask for [help](https://openhands-ai.slack.com/archives/C078L0FUGUX), [feedback](https://openhands-ai.slack.com/archives/C086ARSNMGA), and even ask for a [PR review](https://openhands-ai.slack.com/archives/C08D8FJ5771).\n\n" +
-                    "🙌 Happy hacking! 🙌\n\n" +
-                    "<!-- auto-comment:good-first-issue -->"
-            });
--- a/.gitignore
+++ b/.gitignore
@@ -121,7 +121,6 @@ celerybeat.pid

 # Environments
 .env
-frontend/.env
 .venv
 env/
 venv/
@@ -161,32 +160,8 @@ cython_debug/
 #  and can be added to the global gitignore or merged into this file.  For a more nuclear
 #  option (not recommended) you can uncomment the following to ignore the entire idea folder.
 .idea/
-
-# VS Code: Ignore all but certain files that specify repo-specific settings.
-# https://stackoverflow.com/questions/32964920/should-i-commit-the-vscode-folder-to-source-control
-.vscode/**/*
-!.vscode/extensions.json
-!.vscode/settings.json
-!.vscode/tasks.json
-
-# VS Code extensions/forks:
+.vscode/
 .cursorignore
-.rooignore
-.clineignore
-.windsurfignore
-.cursorrules
-.roorules
-.clinerules
-.windsurfrules
-.cursor/rules
-.roo/rules
-.cline/rules
-.windsurf/rules
-.repomix
-repomix-output.txt
-
-# Emacs backup
-*~

 # evaluation
 evaluation/evaluation_outputs
@@ -194,19 +169,11 @@ evaluation/outputs
 evaluation/swe_bench/eval_workspace*
 evaluation/SWE-bench/data
 evaluation/webarena/scripts/webarena_env.sh
-evaluation/bird/data
-evaluation/gaia/data
-evaluation/gorilla/data
-evaluation/toolqa/data
-evaluation/scienceagentbench/benchmark
-evaluation/commit0_bench/repos
-
-# openhands resolver
-output/

 # frontend

 # dependencies
+frontend/node_modules
 frontend/.pnp
 frontend/bun.lockb
 frontend/yarn.lock
@@ -234,8 +201,6 @@ yarn-error.log*

 logs

-ralph/
-
 # agent
 .envrc
 /workspace
@@ -245,18 +210,10 @@ cache

 # configuration
 config.toml
-config.toml_
 config.toml.bak

+containers/agnostic_sandbox
+
 # swe-bench-eval
 image_build_logs
 run_instance_logs
-
-runtime_*.tar
-
-**/node_modules/
-
-# test results
-test-results
-.sessions
-.eval_sessions
--- a/.nvmrc
+++ b/.nvmrc
@@ -1 +0,0 @@
-22
--- a/.openhands/microagents/documentation.md
+++ b/.openhands/microagents/documentation.md
@@ -1,33 +0,0 @@
---
-name: documentation
-type: knowledge
-version: 1.0.0
-agent: CodeActAgent
-triggers:
- documentation
- docs
- document
---
-
-# Documentation Guidelines
-
-All documentation must be grounded in fact, so you must not make anything up without proper evidence. When you have finished writing documentation, convey to the user what reference source, including web pages, source code, or other sources of documentation you referenced when writing each new fact in the documentation. If you cannot reference a source for anything do not include it in the pull request.
-
-## Best Practices for Documentation
-
-1. **Be Factual**: Only include information that can be verified from reliable sources.
-2. **Cite Sources**: Always reference the source of information (code, web pages, official documentation).
-3. **Be Clear and Concise**: Use simple language and avoid unnecessary jargon.
-4. **Use Examples**: Include practical examples to illustrate concepts.
-5. **Structure Properly**: Use headings, lists, and code blocks to organize information.
-6. **Keep Updated**: Ensure documentation reflects the current state of the code or system.
-
-## Documentation Process
-
-1. Research and gather information from reliable sources
-2. Draft documentation based on verified facts
-3. Review for accuracy and completeness
-4. Include references for all factual statements
-5. Submit only when all information is properly sourced
-
-Remember: If you cannot verify a piece of information, it's better to exclude it than to include potentially incorrect information.
--- a/.openhands/microagents/glossary.md
+++ b/.openhands/microagents/glossary.md
@@ -1,172 +0,0 @@
-# OpenHands Glossary
-
-### Agent
-The core AI entity in OpenHands that can perform software development tasks by interacting with tools, browsing the web, and modifying code.
-
-#### Agent Controller
-A component that manages the agent's lifecycle, handles its state, and coordinates interactions between the agent and various tools.
-
-#### Agent Delegation
-The ability of an agent to hand off specific tasks to other specialized agents for better task completion.
-
-#### Agent Hub
-A central registry of different agent types and their capabilities, allowing for easy agent selection and instantiation.
-
-#### Agent Skill
-A specific capability or function that an agent can perform, such as file manipulation, web browsing, or code editing.
-
-#### Agent State
-The current context and status of an agent, including its memory, active tools, and ongoing tasks.
-
-#### CodeAct Agent
-[A generalist agent in OpenHands](https://arxiv.org/abs/2407.16741) designed to perform tasks by editing and executing code.
-
-### Browser
-A system for web-based interactions and tasks.
-
-#### Browser Gym
-A testing and evaluation environment for browser-based agent interactions and tasks.
-
-#### Web Browser Tool
-A tool that enables agents to interact with web pages and perform web-based tasks.
-
-### Commands
-Terminal and execution related functionality.
-
-#### Bash Session
-A persistent terminal session that maintains state and history for bash command execution.
-This uses tmux under the hood.
-
-### Configuration
-System-wide settings and options.
-
-#### Agent Configuration
-Settings that define an agent's behavior, capabilities, and limitations, including available tools and runtime settings.
-
-#### Configuration Options
-Settings that control various aspects of OpenHands behavior, including runtime, security, and agent settings.
-
-#### LLM Config
-Configuration settings for language models used by agents, including model selection and parameters.
-
-#### LLM Draft Config
-Settings for draft mode operations with language models, typically used for faster, lower-quality responses.
-
-#### Runtime Configuration
-Settings that define how the runtime environment should be set up and operated.
-
-#### Security Options
-Configuration settings that control security features and restrictions.
-
-### Conversation
-A sequence of interactions between a user and an agent, including messages, actions, and their results.
-
-#### Conversation Info
-Metadata about a conversation, including its status, participants, and timeline.
-
-#### Conversation Manager
-A component that handles the creation, storage, and retrieval of conversations.
-
-#### Conversation Metadata
-Additional information about conversations, such as tags, timestamps, and related resources.
-
-#### Conversation Status
-The current state of a conversation, including whether it's active, completed, or failed.
-
-#### Conversation Store
-A storage system for maintaining conversation history and related data.
-
-### Events
-
-#### Event
-Every Conversation comprises a series of Events. Each Event is either an Action or an Observation.
-
-#### Event Stream
-A continuous flow of events that represents the ongoing activities and interactions in the system.
-
-#### Action
-A specific operation or command that an agent executes through available tools, such as running a command or editing a file.
-
-#### Observation
-The response or result returned by a tool after an agent's action, providing feedback about the action's outcome.
-
-### Interface
-Different ways to interact with OpenHands.
-
-#### CLI Mode
-A command-line interface mode for interacting with OpenHands agents without a graphical interface.
-
-#### GUI Mode
-A graphical user interface mode for interacting with OpenHands agents through a web interface.
-
-#### Headless Mode
-A mode of operation where OpenHands runs without a user interface, suitable for automation and scripting.
-
-### Agent Memory
-The system that decides which parts of the Event Stream (i.e. the conversation history) should be passed into each LLM prompt.
-
-#### Memory Store
-A storage system for maintaining agent memory and context across sessions.
-
-#### Condenser
-A component that processes and summarizes conversation history to maintain context while staying within token limits.
-
-#### Truncation
-A very simple Condenser strategy. Reduces conversation history or content to stay within token limits.
-
-### Microagent
-A specialized prompt that enhances OpenHands with domain-specific knowledge, repository-specific context, and task-specific workflows.
-
-#### Microagent Registry
-A central repository of available microagents and their configurations.
-
-#### Public Microagent
-A general-purpose microagent available to all OpenHands users, triggered by specific keywords. Located in `microagents/`.
-
-#### Repository Microagent
-A type of microagent that provides repository-specific context and guidelines, stored in the `.openhands/microagents/` directory.
-
-### Prompt
-Components for managing and processing prompts.
-
-#### Prompt Caching
-A system for caching and reusing common prompts to improve performance.
-
-#### Prompt Manager
-A component that handles the loading, processing, and management of prompts used by agents, including microagents.
-
-#### Response Parsing
-The process of interpreting and structuring responses from language models and tools.
-
-### Runtime
-The execution environment where agents perform their tasks, which can be local, remote, or containerized.
-
-#### Action Execution Server
-A REST API that receives agent actions (e.g. bash commands, python code, browsing actions), executes them in the runtime environment, and returns the results.
-
-#### Action Execution Client
-A component that handles the execution of actions in the runtime environment, managing the communication between the agent and the runtime.
-
-#### Docker Runtime
-A containerized runtime environment that provides isolation and reproducibility for agent operations.
-
-#### E2B Runtime
-A specialized runtime environment built on E2B for secure and isolated code execution.
-
-#### Local Runtime
-A runtime environment that executes on the local machine, suitable for development and testing.
-
-#### Modal Runtime
-A runtime environment built on Modal for scalable and distributed agent operations.
-
-#### Remote Runtime
-A sandboxed environment that executes code and commands remotely, providing isolation and security for agent operations.
-
-#### Runtime Builder
-A component that builds a Docker image for the Action Execution Server based on a user-specified base image.
-
-### Security
-Security-related components and features.
-
-#### Security Analyzer
-A component that checks agent actions for potential security risks.
--- a/.openhands/pre-commit.sh
+++ b/.openhands/pre-commit.sh
@@ -1,124 +0,0 @@
-#!/bin/bash
-
-echo "Running OpenHands pre-commit hook..."
-echo "This hook runs selective linting based on changed files."
-
-# Store the exit code to return at the end
-# This allows us to be additive to existing pre-commit hooks
-EXIT_CODE=0
-
-# Get the list of staged files
-STAGED_FILES=$(git diff --cached --name-only)
-
-# Check if any files match specific patterns
-has_frontend_changes=false
-has_backend_changes=false
-
-# Check each file individually to avoid issues with grep
-for file in $STAGED_FILES; do
-    if [[ $file == frontend/* ]]; then
-        has_frontend_changes=true
-    elif [[ $file == openhands/* || $file == evaluation/* || $file == tests/* ]]; then
-        has_backend_changes=true
-    fi
-done
-
-echo "Analyzing changes..."
-echo "- Frontend changes: $has_frontend_changes"
-echo "- Backend changes: $has_backend_changes"
-
-# Run frontend linting if needed
-if [ "$has_frontend_changes" = true ]; then
-    # Check if we're in a CI environment or if frontend dependencies are missing
-    if [ -n "$CI" ] || ! command -v react-router &> /dev/null || ! command -v vitest &> /dev/null; then
-        echo "Skipping frontend checks (CI environment or missing dependencies detected)."
-        echo "WARNING: Frontend files have changed but frontend checks are being skipped."
-        echo "Please run 'make lint-frontend' manually before submitting your PR."
-    else
-        echo "Running frontend linting..."
-        make lint-frontend
-        if [ $? -ne 0 ]; then
-            echo "Frontend linting failed. Please fix the issues before committing."
-            EXIT_CODE=1
-        else
-            echo "Frontend linting checks passed!"
-        fi
-
-        # Run additional frontend checks
-        if [ -d "frontend" ]; then
-            echo "Running additional frontend checks..."
-            cd frontend || exit 1
-
-            # Run build
-            echo "Running npm build..."
-            npm run build
-            if [ $? -ne 0 ]; then
-                echo "Frontend build failed. Please fix the issues before committing."
-                EXIT_CODE=1
-            fi
-
-            # Run tests
-            echo "Running npm test..."
-            npm test
-            if [ $? -ne 0 ]; then
-                echo "Frontend tests failed. Please fix the failing tests before committing."
-                EXIT_CODE=1
-            fi
-
-            cd ..
-        fi
-    fi
-else
-    echo "Skipping frontend checks (no frontend changes detected)."
-fi
-
-# Run backend linting if needed
-if [ "$has_backend_changes" = true ]; then
-    echo "Running backend linting..."
-    make lint-backend
-    if [ $? -ne 0 ]; then
-        echo "Backend linting failed. Please fix the issues before committing."
-        EXIT_CODE=1
-    else
-        echo "Backend linting checks passed!"
-    fi
-else
-    echo "Skipping backend checks (no backend changes detected)."
-fi
-
-
-# If no specific code changes detected, run basic checks
-if [ "$has_frontend_changes" = false ] && [ "$has_backend_changes" = false ]; then
-    echo "No specific code changes detected. Running basic checks..."
-    if [ -n "$STAGED_FILES" ]; then
-        # Run only basic pre-commit hooks for non-code files
-        poetry run pre-commit run --files $(echo "$STAGED_FILES" | tr '\n' ' ') --hook-stage commit --config ./dev_config/python/.pre-commit-config.yaml
-        if [ $? -ne 0 ]; then
-            echo "Basic checks failed. Please fix the issues before committing."
-            EXIT_CODE=1
-        else
-            echo "Basic checks passed!"
-        fi
-    else
-        echo "No files changed. Skipping basic checks."
-    fi
-fi
-
-# Run any existing pre-commit hooks that might have been installed by the user
-# This makes our hook additive rather than replacing existing hooks
-if [ -f ".git/hooks/pre-commit.local" ]; then
-    echo "Running existing pre-commit hooks..."
-    bash .git/hooks/pre-commit.local
-    if [ $? -ne 0 ]; then
-        echo "Existing pre-commit hooks failed."
-        EXIT_CODE=1
-    fi
-fi
-
-if [ $EXIT_CODE -eq 0 ]; then
-    echo "All pre-commit checks passed!"
-else
-    echo "Some pre-commit checks failed. Please fix the issues before committing."
-fi
-
-exit $EXIT_CODE
--- a/.openhands/setup.sh
+++ b/.openhands/setup.sh
@@ -1,13 +0,0 @@
-#! /bin/bash
-
-echo "Setting up the environment..."
-
-# Install pre-commit package
-python -m pip install pre-commit
-
-# Install pre-commit hooks if .git directory exists
-if [ -d ".git" ]; then
-    echo "Installing pre-commit hooks..."
-    pre-commit install
-    make install-pre-commit-hooks
-fi
--- a/.vscode/settings.json
+++ b/.vscode/settings.json
@@ -1,22 +0,0 @@
-{
-    // force *nix line endings so files don't look modified in container run from Windows clone
-    "files.eol": "\n",
-    "files.trimTrailingWhitespace": true,
-    "files.insertFinalNewline": true,
-
-    "python.defaultInterpreterPath": "./.venv/bin/python",
-    "python.terminal.activateEnvironment": true,
-    "python.analysis.autoImportCompletions": true,
-    "python.analysis.autoSearchPaths": true,
-    "python.analysis.extraPaths": [
-        "./.venv/lib/python3.12/site-packages"
-    ],
-    "python.analysis.packageIndexDepths": [
-        {
-            "name": "openhands",
-            "depth": 10,
-            "includeAllSymbols": true
-        }
-    ],
-    "python.analysis.stubPath": "./.venv/lib/python3.12/site-packages",
-}
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -1,481 +0,0 @@
-This repository contains the code for OpenHands, an automated AI software engineer. It has a Python backend
-(in the `openhands` directory) and React frontend (in the `frontend` directory).
-
-## General Setup:
-To set up the entire repo, including frontend and backend, run `make build`.
-You don't need to do this unless the user asks you to, or if you're trying to run the entire application.
-
-## Running OpenHands with OpenHands:
-To run the full application to debug issues:
-```bash
-export INSTALL_DOCKER=0
-export RUNTIME=local
-make build && make run FRONTEND_PORT=12000 FRONTEND_HOST=0.0.0.0 BACKEND_HOST=0.0.0.0 &> /tmp/openhands-log.txt &
-```
-
-Local run troubleshooting notes:
- If the backend fails with `nc: command not found`, install `netcat-openbsd`.
- If local runtime startup fails with `duplicate session: test-session`, clear the stale tmux session on the default socket: `tmux -S /tmp/tmux-$(id -u)/default kill-session -t test-session`.
- Local runtime browser startup expects Playwright browsers under `~/.cache/playwright`; if needed run `PLAYWRIGHT_BROWSERS_PATH=$HOME/.cache/playwright poetry run playwright install chromium`.
- In this sandbox environment, an inherited `SESSION_API_KEY` can make `/api/v1/settings` return 401 in the browser. Unset it before `make run` when you want to use the local web UI directly.
- In this sandbox, `frontend`'s `npm run dev:mock` / `dev:mock:saas` can start but still be awkward to browse through the work-host proxy. For PR QA screenshots, a reliable fallback is to `npm run build` with the desired `VITE_MOCK_*` env, then serve `build/` with a tiny custom HTTP server that returns the minimal mock JSON endpoints needed by the settings page.
-
-
-IMPORTANT: Before making any changes to the codebase, ALWAYS run `make install-pre-commit-hooks` to ensure pre-commit hooks are properly installed.
-
-Before pushing any changes, you MUST ensure that any lint errors or simple test errors have been fixed.
-
-* If you've made changes to the backend, you should run `pre-commit run --config ./dev_config/python/.pre-commit-config.yaml` (this will run on staged files).
-* If you've made changes to the frontend, you should run `cd frontend && npm run lint:fix && npm run build ; cd ..`
-* If you've made changes to the VSCode extension, you should run `cd openhands/app_server/integrations/vscode && npm run lint:fix && npm run compile ; cd ../../..`
-
-The pre-commit hooks MUST pass successfully before pushing any changes to the repository. This is a mandatory requirement to maintain code quality and consistency.
-
-If either command fails, it may have automatically fixed some issues. You should fix any issues that weren't automatically fixed,
-then re-run the command to ensure it passes. Common issues include:
- Mypy type errors
- Ruff formatting issues
- Trailing whitespace
- Missing newlines at end of files
-
-## Git Best Practices
-
- Prefer specific `git add <filename>` instead of `git add .` to avoid accidentally staging unintended files
- Be especially careful with `git reset --hard` after staging files, as it will remove accidentally staged files
- When remote has new changes, use `git fetch upstream && git rebase upstream/<branch>` on the same branch
-
-## Lockfile Regeneration (Preserve Original Tool Versions)
-
-When regenerating lockfiles (poetry.lock, uv.lock, etc.), you MUST use the same tool version that originally generated the lockfile to avoid unnecessary diff noise. Each lockfile contains a version header indicating which tool version was used.
-
-### Poetry (poetry.lock)
-
-1. Extract the version from the lockfile header:
-   ```bash
-   POETRY_VERSION=$(grep -m1 "^# This file is automatically @generated by Poetry" poetry.lock | sed 's/.*Poetry \([0-9.]*\).*/\1/')
-   ```
-2. If a version is found, install that specific version:
-   ```bash
-   pipx install poetry==$POETRY_VERSION --force
-   ```
-3. Then regenerate the lockfile:
-   ```bash
-   poetry lock --no-update
-   ```
-
-### uv (uv.lock)
-
-1. Extract the version from the lockfile header:
-   ```bash
-   UV_VERSION=$(grep -m1 "^# This file was autogenerated by uv" uv.lock | sed 's/.*uv version \([0-9.]*\).*/\1/')
-   ```
-2. If a version is found, install that specific version:
-   ```bash
-   pipx install uv==$UV_VERSION --force
-   ```
-3. Then regenerate the lockfile:
-   ```bash
-   uv lock
-   ```
-
-This ensures that lockfile updates only contain actual dependency changes, not tool version migration artifacts.
-
-## PR-Specific Artifacts (`.pr/` directory)
-
-When working on a PR that requires design documents, scripts meant for development-only, or other temporary artifacts that should NOT be merged to main, store them in a `.pr/` directory at the repository root.
-
-### Usage
-
-```
-.pr/
-├── design.md       # Design decisions and architecture notes
-├── analysis.md     # Investigation or debugging notes
-├── logs/           # Test output or CI logs for reviewer reference
-└── notes.md        # Any other PR-specific content
-```
-
-### How It Works
-
-1. **Notification**: When `.pr/` exists, a comment is posted to the PR conversation alerting reviewers
-2. **Auto-cleanup**: When the PR is approved, the `.pr/` directory is automatically removed via `.github/workflows/pr-artifacts.yml`
-3. **Fork PRs**: Auto-cleanup cannot push to forks, so manual removal is required before merging
-
-### Important Notes
-
- Do NOT put anything in `.pr/` that needs to be preserved after merge
- The `.pr/` check passes (green ✅) during development — it only posts a notification, not a blocking error
- For fork PRs: You must manually remove `.pr/` before the PR can be merged
-
-### When to Use
-
- Complex refactoring that benefits from written design rationale
- Debugging sessions where you want to document your investigation
- E2E test results or logs that demonstrate a cross-repo feature works
- Feature implementations that need temporary planning docs
- Any analysis that helps reviewers understand the PR but isn't needed long-term
-
-## Repository Structure
-Backend:
- Located in the `openhands` directory
- The current V1 application server lives in `openhands/app_server/`. `make start-backend` still launches `openhands.server.listen:app`, which includes the V1 routes by default unless `ENABLE_V1=0`.
- For V1 web-app docs, LLM setup should point users to the Settings UI.
- Testing:
-  - All tests are in `tests/unit/test_*.py`
-  - To test new code, run `poetry run pytest tests/unit/test_xxx.py` where `xxx` is the appropriate file for the current functionality
-  - Write all tests with pytest
-
-Frontend:
- Located in the `frontend` directory
- Prerequisites: A recent version of NodeJS / NPM
- Setup: Run `npm install` in the frontend directory
- Testing:
-  - Run tests: `npm run test`
-  - To run specific tests: `npm run test -- -t "TestName"`
-  - Our test framework is vitest
- Building:
-  - Build for production: `npm run build`
- Environment Variables:
-  - Set in `frontend/.env` or as environment variables
-  - Available variables: VITE_BACKEND_HOST, VITE_USE_TLS, VITE_INSECURE_SKIP_VERIFY, VITE_FRONTEND_PORT
- Internationalization:
-  - Generate i18n declaration file: `npm run make-i18n`
- Data Fetching & Cache Management:
-  - We use TanStack Query (fka React Query) for data fetching and cache management
-  - Data Access Layer: API client methods are located in `frontend/src/api` and should never be called directly from UI components - they must always be wrapped with TanStack Query
-  - Custom hooks are located in `frontend/src/hooks/query/` and `frontend/src/hooks/mutation/`
-  - Query hooks should follow the pattern use[Resource] (e.g., `useConversationSkills`)
-  - Mutation hooks should follow the pattern use[Action] (e.g., `useDeleteConversation`)
-  - Architecture rule: UI components → TanStack Query hooks → Data Access Layer (`frontend/src/api`) → API endpoints
-  - For SaaS organization management screens, prefer deriving the selected organization from `useOrganizations()` plus the selected org ID store instead of adding a dedicated single-org fetch when only list-level fields (for example `name`) are needed.
-
-
-VSCode Extension:
- Located in the `openhands/app_server/integrations/vscode` directory
- Setup: Run `npm install` in the extension directory
- Linting:
-  - Run linting with fixes: `npm run lint:fix`
-  - Check only: `npm run lint`
-  - Type checking: `npm run typecheck`
- Building:
-  - Compile TypeScript: `npm run compile`
-  - Package extension: `npm run package-vsix`
- Testing:
-  - Run tests: `npm run test`
- Development Best Practices:
-  - Use `vscode.window.createOutputChannel()` for debug logging instead of `showErrorMessage()` popups
-  - Pre-commit process runs both frontend and backend checks when committing extension changes
-
-## Enterprise Directory
-
-The `enterprise/` directory contains additional functionality that extends the open-source OpenHands codebase. This includes:
- Authentication and user management (Keycloak integration)
- Database migrations (Alembic)
- Integration services (GitHub, GitLab, Jira, Linear, Slack)
- Billing and subscription management (Stripe)
- Telemetry and analytics (PostHog, custom metrics framework)
-
-### Enterprise Development Setup
-
-**Prerequisites:**
- Python 3.12
- Poetry (for dependency management)
- Node.js 22.x (for frontend)
- Docker (optional)
-
-**Setup Steps:**
-1. First, build the main OpenHands project: `make build`
-2. Then install enterprise dependencies: `cd enterprise && poetry install --with dev,test` (This can take a very long time. Be patient.)
-3. Set up enterprise pre-commit hooks: `poetry run pre-commit install --config ./dev_config/python/.pre-commit-config.yaml`
-
-**Running Enterprise Tests:**
-```bash
-# Enterprise unit tests (full suite)
-PYTHONPATH=".:$PYTHONPATH" poetry run --project=enterprise pytest --forked -n auto -s -p no:ddtrace -p no:ddtrace.pytest_bdd -p no:ddtrace.pytest_benchmark ./enterprise/tests/unit --cov=enterprise --cov-branch
-
-# Test specific modules (faster for development)
-cd enterprise
-PYTHONPATH=".:$PYTHONPATH" poetry run pytest tests/unit/telemetry/ --confcutdir=tests/unit/telemetry
-
-# Enterprise linting (IMPORTANT: use --show-diff-on-failure to match GitHub CI)
-poetry run pre-commit run --all-files --show-diff-on-failure --config ./dev_config/python/.pre-commit-config.yaml
-```
-
-**Running Enterprise Server:**
-```bash
-cd enterprise
-make start-backend  # Development mode with hot reload
-# or
-make run  # Full application (backend + frontend)
-```
-
-**Key Configuration Files:**
- `enterprise/pyproject.toml` - Enterprise-specific dependencies
- `enterprise/Makefile` - Enterprise build and run commands
- `enterprise/dev_config/python/` - Linting and type checking configuration
- `enterprise/migrations/` - Database migration files
-
-**Database Migrations:**
-Enterprise uses Alembic for database migrations. When making schema changes:
-1. Create migration files in `enterprise/migrations/versions/`
-2. Test migrations thoroughly
-3. The CI will check for migration conflicts on PRs
-
-**Integration Development:**
-The enterprise codebase includes integrations for:
- **GitHub** - PR management, webhooks, app installations
- **GitLab** - Similar to GitHub but for GitLab instances
- **Jira** - Issue tracking and project management
- **Linear** - Modern issue tracking
- **Slack** - Team communication and notifications
-
-Each integration follows a consistent pattern with service classes, storage models, and API endpoints.
-
-**Important Notes:**
- Enterprise code is licensed under Polyform Free Trial License (30-day limit)
- The enterprise server extends the OpenHands server through dynamic imports
- Database changes require careful migration planning in `enterprise/migrations/`
- Always test changes in both OpenHands and enterprise contexts
- Use the enterprise-specific Makefile commands for development
- When the `openhands-ai` package (root project) version has been updated, run `poetry lock` in the `enterprise/` folder to update the version in the enterprise poetry lockfile.
-
-**Enterprise Testing Best Practices:**
-
-**Database Testing:**
- Use SQLite in-memory databases (`sqlite:///:memory:`) for unit tests instead of real PostgreSQL
- Create module-specific `conftest.py` files with database fixtures
- Mock external database connections in unit tests to avoid dependency on running services
- Use real database connections only for integration tests
-
-**Import Patterns:**
- Use relative imports without `enterprise.` prefix in enterprise code
- Example: `from storage.database import a_session_maker` not `from enterprise.storage.database import a_session_maker`
- This ensures code works in both OpenHands and enterprise contexts
-
-**Test Structure:**
- Place tests in `enterprise/tests/unit/` following the same structure as the source code
- Use `--confcutdir=tests/unit/[module]` when testing specific modules
- Create comprehensive fixtures for complex objects (databases, external services)
- Write platform-agnostic tests (avoid hardcoded OS-specific assertions)
-
-**Mocking Strategy:**
- Use `AsyncMock` for async operations and `MagicMock` for complex objects
- Mock all external dependencies (databases, APIs, file systems) in unit tests
- Use `patch` with correct import paths (e.g., `telemetry.registry.logger` not `enterprise.telemetry.registry.logger`)
- Test both success and failure scenarios with proper error handling
-
-**Coverage Goals:**
- Aim for 90%+ test coverage on new enterprise modules
- Focus on critical business logic and error handling paths
- Use `--cov-report=term-missing` to identify uncovered lines
-
-**Troubleshooting:**
- If tests fail, ensure all dependencies are installed: `poetry install --with dev,test`
- For database issues, check migration status and run migrations if needed
- For frontend issues, ensure the main OpenHands frontend is built: `make build`
- Check logs in the `logs/` directory for runtime issues
- If tests fail with import errors, verify `PYTHONPATH=".:$PYTHONPATH"` is set
- **If GitHub CI fails but local linting passes**: Always use `--show-diff-on-failure` flag to match CI behavior exactly
-
-## Template for Github Pull Request
-
-If you are starting a pull request (PR), please follow the template in `.github/pull_request_template.md`.
-
-## Implementation Details
-
-These details may or may not be useful for your current task.
-
-### Conversation State Management
-
-#### Agent State and Sandbox Status:
-The frontend uses `useAgentState` hook (`frontend/src/hooks/use-agent-state.ts`) to determine the current conversation state. This hook:
- Returns `curAgentState` (AgentState enum) for UI state determination
- Returns `isArchived` flag when `sandbox_status === "MISSING"` (archived conversations)
- Prioritizes live WebSocket execution status over cached API data
-
-#### Archived Conversations (sandbox_status === "MISSING"):
-When a conversation's sandbox is no longer available (archived):
- `useAgentState` returns `AgentState.STOPPED` and `isArchived: true`
- Chat input is replaced with an archived banner (`ArchivedBanner` component)
- VS Code tab, Terminal, and Planner show read-only messages instead of loading states
- All interactive elements that require a running sandbox are disabled
-
-#### Testing useAgentState:
-When mocking `useAgentState` in tests, always include the `isArchived` property:
-```typescript
-vi.mock("#/hooks/use-agent-state", () => ({
-  useAgentState: () => ({
-    curAgentState: AgentState.AWAITING_USER_INPUT,
-    isArchived: false,
-  }),
-}));
-```
-
-### Microagents
-
-Microagents are specialized prompts that enhance OpenHands with domain-specific knowledge and task-specific workflows. They are Markdown files that can include frontmatter for configuration.
-
-#### Types:
- **Public Microagents**: Located in `microagents/`, available to all users
- **Repository Microagents**: Located in `.openhands/microagents/`, specific to this repository
-
-#### Loading Behavior:
- **Without frontmatter**: Always loaded into LLM context
- **With triggers in frontmatter**: Only loaded when user's message matches the specified trigger keywords
-
-#### Structure:
-```yaml
---
-triggers:
- keyword1
- keyword2
---
-# Microagent Content
-Your specialized knowledge and instructions here...
-```
-
-### Frontend
-
-#### Action Handling:
- Actions are defined in `frontend/src/types/action-type.ts`
- The `HANDLED_ACTIONS` array in `frontend/src/state/chat-slice.ts` determines which actions are displayed as collapsible UI elements
- To add a new action type to the UI:
-  1. Add the action type to the `HANDLED_ACTIONS` array
-  2. Implement the action handling in `addAssistantAction` function in chat-slice.ts
-  3. Add a translation key in the format `ACTION_MESSAGE$ACTION_NAME` to the i18n files
- Actions with `thought` property are displayed in the UI based on their action type:
-  - Regular actions (like "run", "edit") display the thought as a separate message
-  - Special actions (like "think") are displayed as collapsible elements only
-
-#### Adding User Settings:
- To add a new user setting to OpenHands, follow these steps:
-  1. Add the setting to the frontend:
-     - Add the setting to the `Settings` type in `frontend/src/types/settings.ts`
-     - Add the setting to the `ApiSettings` type in the same file
-     - Add the setting with an appropriate default value to `DEFAULT_SETTINGS` in `frontend/src/services/settings.ts`
-     - Update the `useSettings` hook in `frontend/src/hooks/query/use-settings.ts` to map the API response
-     - Update the `useSaveSettings` hook in `frontend/src/hooks/mutation/use-save-settings.ts` to include the setting in API requests
-     - Add UI components (like toggle switches) in the appropriate settings screen (e.g., `frontend/src/routes/app-settings.tsx`)
-     - Add i18n translations for the setting name and any tooltips in `frontend/src/i18n/translation.json`
-     - Add the translation key to `frontend/src/i18n/declaration.ts`
-  2. Add the setting to the backend:
-     - Add the setting to the `Settings` model in `openhands/storage/data_models/settings.py`
-     - Update any relevant backend code to apply the setting (e.g., in session creation)
-
-#### Settings UI Patterns:
-
-There are two main patterns for saving settings in the OpenHands frontend:
-
-**Pattern 1: Entity-based Resources (Immediate Save)**
- Used for: API Keys, Secrets, MCP Servers
- Behavior: Changes are saved immediately when user performs actions (add/edit/delete)
- Implementation:
-  - No "Save Changes" button
-  - No local state management or `isDirty` tracking
-  - Uses dedicated mutation hooks for each operation (e.g., `use-add-mcp-server.ts`, `use-delete-mcp-server.ts`)
-  - Each mutation triggers immediate API call with query invalidation for UI updates
-  - Example: MCP settings, API Keys & Secrets tabs
- Benefits: Simpler UX, no risk of losing changes, consistent with modern web app patterns
-
-**Pattern 2: Form-based Settings (Manual Save)**
- Used for: Application settings, LLM configuration
- Behavior: Changes are accumulated locally and saved when user clicks "Save Changes"
- Implementation:
-  - Has "Save Changes" button that becomes enabled when changes are detected
-  - Uses local state management with `isDirty` tracking
-  - Uses `useSaveSettings` hook to save all changes at once
-  - Example: LLM tab, Application tab
- Benefits: Allows bulk changes, explicit save action, can validate all fields before saving
-
-**When to use each pattern:**
- Use Pattern 1 (Immediate Save) for entity management where each item is independent
- Use Pattern 2 (Manual Save) for configuration forms where settings are interdependent or need validation
- Git provider tokens in the local/OSS integrations settings are managed through the V1 secrets endpoints (`POST`/`DELETE /api/v1/secrets/git-providers`). Do not reuse the logout flow for disconnecting tokens; `useLogout` is for actual app logout and still targets legacy OSS logout behavior.
-
-### Adding New LLM Models
-
-To add a new LLM model to OpenHands, you need to update multiple files across both frontend and backend:
-
-#### Model Configuration Procedure:
-
-1. **Frontend Model Arrays** (`frontend/src/utils/verified-models.ts`):
-   - Add the model to `VERIFIED_MODELS` array (main list of all verified models)
-   - Add to provider-specific arrays based on the model's provider:
-     - `VERIFIED_OPENAI_MODELS` for OpenAI models
-     - `VERIFIED_ANTHROPIC_MODELS` for Anthropic models
-     - `VERIFIED_MISTRAL_MODELS` for Mistral models
-     - `VERIFIED_OPENHANDS_MODELS` for models available through OpenHands provider
-
-2. **Backend CLI Integration** (`openhands/cli/utils.py`):
-   - Add the model to the appropriate `VERIFIED_*_MODELS` arrays
-   - This ensures the model appears in CLI model selection
-
-3. **Backend Model List** (`openhands/utils/llm.py`):
-   - **CRITICAL**: Add the model to the `openhands_models` list (lines 57-66) if using OpenHands provider
-   - This is required for the model to appear in the frontend model selector
-   - Format: `'openhands/model-name'` (e.g., `'openhands/o3'`)
-
-4. **Backend LLM Configuration** (`openhands/llm/llm.py`):
-   - Add to feature-specific arrays based on model capabilities:
-     - `FUNCTION_CALLING_SUPPORTED_MODELS` if the model supports function calling
-     - `REASONING_EFFORT_SUPPORTED_MODELS` if the model supports reasoning effort parameters
-     - `CACHE_PROMPT_SUPPORTED_MODELS` if the model supports prompt caching
-     - `MODELS_WITHOUT_STOP_WORDS` if the model doesn't support stop words
-
-5. **Validation**:
-   - Run backend linting: `pre-commit run --config ./dev_config/python/.pre-commit-config.yaml`
-   - Run frontend linting: `cd frontend && npm run lint:fix`
-   - Run frontend build: `cd frontend && npm run build`
-
-#### Model Verification Arrays:
-
- **VERIFIED_MODELS**: Main array of all verified models shown in the UI
- **VERIFIED_OPENAI_MODELS**: OpenAI models (LiteLLM doesn't return provider prefix)
- **VERIFIED_ANTHROPIC_MODELS**: Anthropic models (LiteLLM doesn't return provider prefix)
- **VERIFIED_MISTRAL_MODELS**: Mistral models (LiteLLM doesn't return provider prefix)
- **VERIFIED_OPENHANDS_MODELS**: Models available through OpenHands managed provider
-
-#### Model Feature Support Arrays:
-
- **FUNCTION_CALLING_SUPPORTED_MODELS**: Models that support structured function calling
- **REASONING_EFFORT_SUPPORTED_MODELS**: Models that support reasoning effort parameters (like o1, o3)
- **CACHE_PROMPT_SUPPORTED_MODELS**: Models that support prompt caching for efficiency
- **MODELS_WITHOUT_STOP_WORDS**: Models that don't support stop word parameters
-
-#### Frontend Model Integration:
-
- Models are automatically available in the model selector UI once added to verified arrays
- The `extractModelAndProvider` utility automatically detects provider from model arrays
- Provider-specific models are grouped and prioritized in the UI selection
-
-#### CLI Model Integration:
-
- Models appear in CLI provider selection based on the verified arrays
- The `organize_models_and_providers` function groups models by provider
- Default model selection prioritizes verified models for each provider
-
-### Sandbox Settings API (SDK Credential Inheritance)
-
-The sandbox settings API allows SDK-created conversations to inherit the user's SaaS credentials
-(LLM config, secrets) securely via `LookupSecret`. Raw secret values only flow SaaS→sandbox,
-never through the SDK client.
-
-#### User Credentials with Exposed Secrets (in `openhands/app_server/user/user_router.py`):
- `GET /api/v1/users/me?expose_secrets=true` → Full user settings with unmasked secrets (e.g., `llm_api_key`)
- `GET /api/v1/users/me` → Full user settings (secrets masked, Bearer only)
-
-Auth requirements for `expose_secrets=true`:
- Bearer token (proves user identity via `OPENHANDS_API_KEY`)
- `X-Session-API-Key` header (proves caller has an active sandbox owned by the authenticated user)
-
-Called by `workspace.get_llm()` in the SDK to retrieve LLM config with the API key.
-
-#### Sandbox-Scoped Secrets Endpoints (in `openhands/app_server/sandbox/sandbox_router.py`):
- `GET /sandboxes/{id}/settings/secrets` → list secret names (no values)
- `GET /sandboxes/{id}/settings/secrets/{name}` → raw secret value (called FROM sandbox)
-
-#### Auth: `X-Session-API-Key` header, validated via `SandboxService.get_sandbox_by_session_api_key()`
-
-#### Related SDK code (in `software-agent-sdk` repo):
- `openhands/sdk/llm/llm.py`: `LLM.api_key` accepts `SecretSource` (including `LookupSecret`)
- `openhands/workspace/cloud/workspace.py`: `get_llm()` and `get_secrets()` return LookupSecret-backed objects
- Tests: `tests/sdk/llm/test_llm_secret_source_api_key.py`, `tests/workspace/test_cloud_workspace_sdk_settings.py`
--- a/CITATION.cff
+++ b/CITATION.cff
@@ -1,55 +0,0 @@
-cff-version: 1.2.0
-message: "If you use this software, please cite it using the following metadata."
-title: "OpenHands: An Open Platform for AI Software Developers as Generalist Agents"
-authors:
-  - family-names: Wang
-    given-names: Xingyao
-  - family-names: Li
-    given-names: Boxuan
-  - family-names: Song
-    given-names: Yufan
-  - family-names: Xu
-    given-names: Frank F.
-  - family-names: Tang
-    given-names: Xiangru
-  - family-names: Zhuge
-    given-names: Mingchen
-  - family-names: Pan
-    given-names: Jiayi
-  - family-names: Song
-    given-names: Yueqi
-  - family-names: Li
-    given-names: Bowen
-  - family-names: Singh
-    given-names: Jaskirat
-  - family-names: Tran
-    given-names: Hoang H.
-  - family-names: Li
-    given-names: Fuqiang
-  - family-names: Ma
-    given-names: Ren
-  - family-names: Zheng
-    given-names: Mingzhang
-  - family-names: Qian
-    given-names: Bill
-  - family-names: Shao
-    given-names: Yanjun
-  - family-names: Muennighoff
-    given-names: Niklas
-  - family-names: Zhang
-    given-names: Yizhe
-  - family-names: Hui
-    given-names: Binyuan
-  - family-names: Lin
-    given-names: Junyang
-  - family-names: Brennan
-    given-names: Robert
-  - family-names: Peng
-    given-names: Hao
-  - family-names: Ji
-    given-names: Heng
-  - family-names: Neubig
-    given-names: Graham
-year: 2024
-doi: "10.48550/arXiv.2407.16741"
-url: "https://arxiv.org/abs/2407.16741"
--- a/1
+++ b/1
@@ -1 +0,0 @@
-docs.all-hands.dev
--- a/CODE_OF_CONDUCT.md
+++ b/CODE_OF_CONDUCT.md
@@ -18,24 +18,24 @@ diverse, inclusive, and healthy community.
 Examples of behavior that contributes to a positive environment for our
 community include:

-* Demonstrating empathy and kindness toward other people.
-* Being respectful of differing opinions, viewpoints, and experiences.
-* Giving and gracefully accepting constructive feedback.
+* Demonstrating empathy and kindness toward other people
+* Being respectful of differing opinions, viewpoints, and experiences
+* Giving and gracefully accepting constructive feedback
 * Accepting responsibility and apologizing to those affected by our mistakes,
-  and learning from the experience.
+  and learning from the experience
 * Focusing on what is best not just for us as individuals, but for the overall
-  community.
+  community

 Examples of unacceptable behavior include:

 * The use of sexualized language or imagery, and sexual attention or advances of
-  any kind.
-* Trolling, insulting or derogatory comments, and personal or political attacks.
-* Public or private harassment.
+  any kind
+* Trolling, insulting or derogatory comments, and personal or political attacks
+* Public or private harassment
 * Publishing others' private information, such as a physical or email address,
-  without their explicit permission.
+  without their explicit permission
 * Other conduct which could reasonably be considered inappropriate in a
-  professional setting.
+  professional setting

 ## Enforcement Responsibilities

@@ -61,7 +61,7 @@ representative at an online or offline event.

 Instances of abusive, harassing, or otherwise unacceptable behavior may be
 reported to the community leaders responsible for enforcement at
-contact@openhands.dev.
+contact@all-hands.dev
 All complaints will be reviewed and investigated promptly and fairly.

 All community leaders are obligated to respect the privacy and security of the
@@ -113,25 +113,6 @@ individual, or aggression toward or disparagement of classes of individuals.
 **Consequence**: A permanent ban from any sort of public interaction within the
 community.

-### Slack Etiquettes
-
-These Slack etiquette guidelines are designed to foster an inclusive, respectful, and productive environment for all
-community members. By following these best practices, we ensure effective communication and collaboration while
-minimizing disruptions. Let’s work together to build a supportive and welcoming community!
-
- Communicate respectfully and professionally, avoiding sarcasm or harsh language, and remember that tone can be difficult to interpret in text.
- Use threads for specific discussions to keep channels organized and easier to follow.
- Tag others only when their input is critical or urgent, and use @here, @channel or @everyone sparingly to minimize disruptions.
- Be patient, as open-source contributors and maintainers often have other commitments and may need time to respond.
- Post questions or discussions in the most relevant channel (e.g., for [slack - #general](https://openhands-ai.slack.com/archives/C06P5NCGSFP) for general topics, [slack - #questions](https://openhands-ai.slack.com/archives/C06U8UTKSAD) for queries/questions.
- When asking for help or raising issues, include necessary details like links, screenshots, or clear explanations to provide context.
- Keep discussions in public channels whenever possible to allow others to benefit from the conversation, unless the matter is sensitive or private.
- Always adhere to [our standards](https://github.com/OpenHands/OpenHands/blob/main/CODE_OF_CONDUCT.md#our-standards) to ensure a welcoming and collaborative environment.
- If you choose to mute a channel, consider setting up alerts for topics that still interest you to stay engaged.
-   For Slack, Go to Settings → Notifications → My Keywords to add specific keywords that will notify you when mentioned.
-   For example, if you're here for discussions about LLMs, mute the channel if it’s too busy, but set notifications to
-   alert you only when “LLMs” appears in messages.
-
 ## Attribution

 This Code of Conduct is adapted from the [Contributor Covenant][homepage],
--- a/COMMUNITY.md
+++ b/COMMUNITY.md
@@ -1,58 +0,0 @@
-# The OpenHands Community
-
-OpenHands is a community of engineers, academics, and enthusiasts reimagining software development for an AI-powered
-world.
-
-## Mission
-
-It’s very clear that AI is changing software development. We want the developer community to drive that change
-organically, through open source.
-
-So we’re not just building friendly interfaces for AI-driven development. We’re publishing _building blocks_ that
-empower developers to create new experiences, tailored to your own habits, needs, and imagination.
-
-## Ethos
-
-We have two core values: **high openness** and **high agency**. While we don’t expect everyone in the community to
-embody these values, we want to establish them as norms.
-
-### High Openness
-
-We welcome anyone and everyone into our community by default. You don’t have to be a software developer to help us
-build. You don’t have to be pro-AI to help us learn.
-
-Our plans, our work, our successes, and our failures are all public record. We want the world to see not just the
-fruits of our work, but the whole process of growing it.
-
-We welcome thoughtful criticism, whether it’s a comment on a PR or feedback on the community as a whole.
-
-### High Agency
-
-Everyone should feel empowered to contribute to OpenHands. Whether it’s by making a PR, hosting an event, sharing
-feedback, or just asking a question, don’t hold back!
-
-OpenHands gives everyone the building blocks to create state-of-the-art developer experiences. We experiment constantly
-and love building new things.
-
-Coding, development practices, and communities are changing rapidly. We won’t hesitate to change direction and make big bets.
-
-## Relationship to All Hands
-
-OpenHands is supported by the for-profit organization [All Hands AI, Inc](https://www.openhands.dev/).
-
-All Hands was founded by three of the first major contributors to OpenHands:
-
- Xingyao Wang, a UIUC PhD candidate who got OpenHands to the top of the SWE-bench leaderboards
- Graham Neubig, a CMU Professor who rallied the academic community around OpenHands
- Robert Brennan, a software engineer who architected the user-facing features of OpenHands
-
-All Hands is an important part of the OpenHands ecosystem. We’ve raised over $20M--mainly to hire developers and
-researchers who can work on OpenHands full-time, and to provide them with expensive infrastructure. ([Join us!](https://allhandsai.applytojob.com/apply/))
-
-But we see OpenHands as much larger, and ultimately more important, than All Hands. When our financial responsibility
-to investors is at odds with our social responsibility to the community—as it inevitably will be, from time to time—we
-promise to navigate that conflict thoughtfully and transparently.
-
-At some point, we may transfer custody of OpenHands to an open source foundation. But for now,
-the [Benevolent Dictator approach](http://www.catb.org/~esr/writings/cathedral-bazaar/homesteading/ar01s16.html) helps us move forward with speed and intention. If we ever forget the
-“benevolent” part, please: fork us.
--- a/CONTRIBUTING.md
+++ b/CONTRIBUTING.md
@@ -1,102 +1,96 @@
 # Contributing

-Thanks for your interest in contributing to OpenHands! We're building the future of AI-powered software development, and we'd love for you to be part of this journey.
+Thanks for your interest in contributing to OpenDevin! We welcome and appreciate contributions. 

-## Our Vision
+## How Can I Contribute?

-The OpenHands community is built around the belief that AI and AI agents are going to fundamentally change the way we build software. If this is true, we should do everything we can to make sure that the benefits provided by such powerful technology are accessible to everyone.
+There are many ways that you can contribute:

-We believe in the power of open source to democratize access to cutting-edge AI technology. Just as the internet transformed how we share information, we envision a world where AI-powered development tools are available to every developer, regardless of their background or resources.
+1. **Download and use** OpenDevin, and send [issues](https://github.com/OpenDevin/OpenDevin/issues) when you encounter something that isn't working or a feature that you'd like to see.
+2. **Send feedback** after each session by [clicking the thumbs-up thumbs-down buttons](https://opendevin.github.io/OpenDevin/modules/usage/feedback), so we can see where things are working and failing, and also build an open dataset for training code agents.
+3. **Improve the Codebase** by sending PRs (see details below). In particular, we have some [good first issue](https://github.com/OpenDevin/OpenDevin/labels/good%20first%20issue) issues that may be ones to start on.

-## Getting Started
+## Understanding OpenDevin's CodeBase

-### Quick Ways to Contribute
+To understand the codebase, please refer to the README in each module:
+- [frontend](./frontend/README.md)
+- [agenthub](./agenthub/README.md)
+- [evaluation](./evaluation/README.md)
+- [opendevin](./opendevin/README.md)
+    - [server](./opendevin/server/README.md)

- **Use OpenHands** and [report issues](https://github.com/OpenHands/OpenHands/issues) you encounter
- **Give feedback** using the thumbs-up/thumbs-down buttons after each session
- **Star our repository** on [GitHub](https://github.com/OpenHands/OpenHands)
- **Share OpenHands** with other developers
+When you write code, it is also good to write tests. Please navigate to the `tests` folder to see existing test suites.
+At the moment, we have two kinds of tests: `unit` and `integration`. Please refer to the README for each test suite. These tests also run on GitHub's continuous integration to ensure quality of the project.

-### Set Up Your Development Environment
+## Sending Pull Requests to OpenDevin

- **Requirements**: Linux/Mac/WSL, Docker, Python 3.12, Node.js 22+, Poetry 1.8+
- **Quick setup**: `make build`
- **Run locally**: `make run`
- **LLM setup (V1 web app)**: configure your model and API key in the Settings UI after the app starts
+### 1. Fork the Official Repository
+Fork the [OpenDevin repository](https://github.com/OpenDevin/OpenDevin) into your own account.
+Clone your own forked repository into your local environment:

-Full details in our [Development Guide](./Development.md).
+```shell
+git clone git@github.com:<YOUR-USERNAME>/OpenDevin.git
+```

-### Find Your First Issue
+### 2. Configure Git

- Browse [good first issues](https://github.com/OpenHands/OpenHands/labels/good%20first%20issue)
- Check our [project boards](https://github.com/OpenHands/OpenHands/projects) for organized tasks
- Join our [Slack community](https://openhands.dev/joinslack) to ask what needs help
+Set the official repository as your [upstream](https://www.atlassian.com/git/tutorials/git-forks-and-upstreams) to synchronize with the latest update in the official repository.
+Add the original repository as upstream:

-## Understanding the Codebase
+```shell
+cd OpenDevin
+git remote add upstream git@github.com:OpenDevin/OpenDevin.git
+```

- **[Frontend](./frontend/README.md)** - React application
- **[App Server (V1)](./openhands/app_server/README.md)** - Current FastAPI application server and REST API modules
- **[Evaluation](https://github.com/OpenHands/benchmarks)** - Testing and benchmarks
+Verify that the remote is set:

-## What Can You Build?
+```shell
+git remote -v
+```

-### Frontend & UI/UX
- React & TypeScript development
- UI/UX improvements
- Mobile responsiveness
- Component libraries
+You should see both `origin` and `upstream` in the output.

-For bigger changes, join the #proj-gui channel in [Slack](https://openhands.dev/joinslack) first.
+### 3. Synchronize with Official Repository
+Synchronize latest commit with official repository before coding:

-### Agent Development
- Prompt engineering
- New agent types
- Agent evaluation
- Multi-agent systems
+```shell
+git fetch upstream
+git checkout main
+git merge upstream/main
+git push origin main
+```

-We use [SWE-bench](https://www.swebench.com/) to evaluate agents.
+### 4. Set up the Development Environment

-### Backend & Infrastructure
- Python development
- Runtime systems (Docker containers, sandboxes)
- Cloud integrations
- Performance optimization
+We have a separate doc [Development.md](https://github.com/OpenDevin/OpenDevin/blob/main/Development.md) that tells you how to set up a development workflow.

-### Testing & Quality Assurance
- Unit testing
- Integration testing
- Bug hunting
- Performance testing
+### 5. Write Code and Commit It

-### Documentation & Education
- Technical documentation
- Translation
- Community support
+Once you have done this, you can write code, test it, and commit it to a branch (replace `my_branch` with an appropriate name):

-## Pull Request Process
+```shell
+git checkout -b my_branch
+git add .
+git commit
+git push origin my_branch
+```

-### Small Improvements
- Quick review and approval
- Ensure CI tests pass
- Include clear description of changes
+### 6. Open a Pull Request

-### Core Agent Changes
-These are evaluated based on:
- **Accuracy** - Does it make the agent better at solving problems?
- **Efficiency** - Does it improve speed or reduce resource usage?
- **Code Quality** - Is the code maintainable and well-tested?
+* On GitHub, go to the page of your forked repository, and create a Pull Request:
+   - Click on `Branches`
+   - Click on the `...` beside your branch and click on `New pull request`
+   - Set `base repository` to `OpenDevin/OpenDevin`
+   - Set `base` to `main`
+   - Click `Create pull request`
+  
+The PR should appear in [OpenDevin PRs](https://github.com/OpenDevin/OpenDevin/pulls).

-Discuss major changes in [GitHub issues](https://github.com/OpenHands/OpenHands/issues) or [Slack](https://openhands.dev/joinslack) first.
+Then the OpenDevin team will review your code.

-## Sending Pull Requests to OpenHands
-
-You'll need to fork our repository to send us a Pull Request. You can learn more
-about how to fork a GitHub repo and open a PR with your changes in [this article](https://medium.com/swlh/forks-and-pull-requests-how-to-contribute-to-github-repos-8843fac34ce8).
-
-You may also check out previous PRs in the [PR list](https://github.com/OpenHands/OpenHands/pulls).
-
-### Pull Request Title Format
+## PR Rules

+### 1. Pull Request title
 As described [here](https://github.com/commitizen/conventional-commit-types/blob/master/index.json), a valid PR title should begin with one of the following prefixes:

 - `feat`: A new feature
@@ -115,27 +109,9 @@ For example, a PR title could be:
 - `refactor: modify package path`
 - `feat(frontend): xxxx`, where `(frontend)` means that this PR mainly focuses on the frontend component.

-### Pull Request Description
+You may also check out previous PRs in the [PR list](https://github.com/OpenDevin/OpenDevin/pulls).

- Explain what the PR does and why
- Link to related issues
- Include screenshots for UI changes
- If your changes are user-facing (e.g. a new feature in the UI, a change in behavior, or a bugfix),
-  please include a short message that we can add to our changelog
+### 2. Pull Request description
+- If your PR is small (such as a typo fix), you can go brief.
+- If it contains a lot of changes, it's better to write more details.

-## Becoming a Maintainer
-
-For contributors who have made significant and sustained contributions to the project, there is a possibility of joining the maintainer team.
-The process for this is as follows:
-
-1. Any contributor who has made sustained and high-quality contributions to the codebase can be nominated by any maintainer. If you feel that you may qualify you can reach out to any of the maintainers that have reviewed your PRs and ask if you can be nominated.
-2. Once a maintainer nominates a new maintainer, there will be a discussion period among the maintainers for at least 3 days.
-3. If no concerns are raised the nomination will be accepted by acclamation, and if concerns are raised there will be a discussion and possible vote.
-
-Note that just making many PRs does not immediately imply that you will become a maintainer. We will be looking at sustained high-quality contributions over a period of time, as well as good teamwork and adherence to our [Code of Conduct](./CODE_OF_CONDUCT.md).
-
-## Need Help?
-
- **Slack**: [Join our community](https://openhands.dev/joinslack)
- **GitHub Issues**: [Open an issue](https://github.com/OpenHands/OpenHands/issues)
- **Email**: contact@openhands.dev
--- a/CREDITS.md
+++ b/CREDITS.md
@@ -1,328 +0,0 @@
-# Credits
-
-## Contributors
-
-We would like to thank all the [contributors](https://github.com/OpenHands/OpenHands/graphs/contributors) who have
-helped make OpenHands possible. We greatly appreciate your dedication and hard work.
-
-## Open Source Projects
-
-OpenHands includes and adapts the following open source projects. We are grateful for their contributions to the
-open source community:
-
-#### [SWE Agent](https://github.com/princeton-nlp/swe-agent)
-   - License: MIT License
-   - Description: Adapted for use in OpenHands's agent hub
-
-#### [Aider](https://github.com/paul-gauthier/aider)
-   - License: Apache License 2.0
-   - Description: AI pair programming tool. OpenHands has adapted and integrated its linter module for code-related tasks.
-
-#### [BrowserGym](https://github.com/ServiceNow/BrowserGym)
-   - License: Apache License 2.0
-   - Description: Adapted in implementing the browsing agent
-
-### Reference Implementations for Evaluation Benchmarks
-
-OpenHands integrates code of the reference implementations for the following agent evaluation benchmarks:
-
-#### [HumanEval](https://github.com/openai/human-eval)
-   - License: MIT License
-
-#### [DSP](https://github.com/microsoft/DataScienceProblems)
-   - License: MIT License
-
-#### [HumanEvalPack](https://github.com/bigcode-project/bigcode-evaluation-harness)
-   - License: Apache License 2.0
-
-#### [AgentBench](https://github.com/THUDM/AgentBench)
-   - License: Apache License 2.0
-
-#### [SWE-Bench](https://github.com/princeton-nlp/SWE-bench)
-   - License: MIT License
-
-#### [BIRD](https://bird-bench.github.io/)
-   - License: MIT License
-   - Dataset: CC-BY-SA 4.0
-
-#### [Gorilla APIBench](https://github.com/ShishirPatil/gorilla)
-   - License: Apache License 2.0
-
-#### [GPQA](https://github.com/idavidrein/gpqa)
-   - License: MIT License
-
-#### [ProntoQA](https://github.com/asaparov/prontoqa)
-   - License: Apache License 2.0
-
-## Open Source licenses
-
-### MIT License
-
-Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated
-documentation files (the "Software"), to deal in the Software without restriction, including without limitation the
-rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit
-persons to whom the Software is furnished to do so, subject to the following conditions:
-
-The above copyright notice and this permission notice shall be included in all copies or substantial portions of the
-Software.
-
-THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO
-THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS
-OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR
-OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
-
-### BSD 3-Clause License
-
-Redistribution and use in source and binary forms, with or without modification, are permitted provided that the
-following conditions are met:
-
-1. Redistributions of source code must retain the above copyright notice, this list of conditions and the following
-   disclaimer.
-
-2. Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following
-   disclaimer in the documentation and/or other materials provided with the distribution.
-
-3. Neither the name of the copyright holder nor the names of its contributors may be used to endorse or promote
-   products derived from this software without specific prior written permission.
-
-THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS" AND ANY EXPRESS OR IMPLIED WARRANTIES,
-INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
-DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL,
-SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
-SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY,
-WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
-OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
-
-### Apache License 2.0
-
-
-                                 Apache License
-                           Version 2.0, January 2004
-                        http://www.apache.org/licenses/
-
-   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
-
-   1. Definitions.
-
-      "License" shall mean the terms and conditions for use, reproduction,
-      and distribution as defined by Sections 1 through 9 of this document.
-
-      "Licensor" shall mean the copyright owner or entity authorized by
-      the copyright owner that is granting the License.
-
-      "Legal Entity" shall mean the union of the acting entity and all
-      other entities that control, are controlled by, or are under common
-      control with that entity. For the purposes of this definition,
-      "control" means (i) the power, direct or indirect, to cause the
-      direction or management of such entity, whether by contract or
-      otherwise, or (ii) ownership of fifty percent (50%) or more of the
-      outstanding shares, or (iii) beneficial ownership of such entity.
-
-      "You" (or "Your") shall mean an individual or Legal Entity
-      exercising permissions granted by this License.
-
-      "Source" form shall mean the preferred form for making modifications,
-      including but not limited to software source code, documentation
-      source, and configuration files.
-
-      "Object" form shall mean any form resulting from mechanical
-      transformation or translation of a Source form, including but
-      not limited to compiled object code, generated documentation,
-      and conversions to other media types.
-
-      "Work" shall mean the work of authorship, whether in Source or
-      Object form, made available under the License, as indicated by a
-      copyright notice that is included in or attached to the work
-      (an example is provided in the Appendix below).
-
-      "Derivative Works" shall mean any work, whether in Source or Object
-      form, that is based on (or derived from) the Work and for which the
-      editorial revisions, annotations, elaborations, or other modifications
-      represent, as a whole, an original work of authorship. For the purposes
-      of this License, Derivative Works shall not include works that remain
-      separable from, or merely link (or bind by name) to the interfaces of,
-      the Work and Derivative Works thereof.
-
-      "Contribution" shall mean any work of authorship, including
-      the original version of the Work and any modifications or additions
-      to that Work or Derivative Works thereof, that is intentionally
-      submitted to Licensor for inclusion in the Work by the copyright owner
-      or by an individual or Legal Entity authorized to submit on behalf of
-      the copyright owner. For the purposes of this definition, "submitted"
-      means any form of electronic, verbal, or written communication sent
-      to the Licensor or its representatives, including but not limited to
-      communication on electronic mailing lists, source code control systems,
-      and issue tracking systems that are managed by, or on behalf of, the
-      Licensor for the purpose of discussing and improving the Work, but
-      excluding communication that is conspicuously marked or otherwise
-      designated in writing by the copyright owner as "Not a Contribution."
-
-      "Contributor" shall mean Licensor and any individual or Legal Entity
-      on behalf of whom a Contribution has been received by Licensor and
-      subsequently incorporated within the Work.
-
-   2. Grant of Copyright License. Subject to the terms and conditions of
-      this License, each Contributor hereby grants to You a perpetual,
-      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
-      copyright license to reproduce, prepare Derivative Works of,
-      publicly display, publicly perform, sublicense, and distribute the
-      Work and such Derivative Works in Source or Object form.
-
-   3. Grant of Patent License. Subject to the terms and conditions of
-      this License, each Contributor hereby grants to You a perpetual,
-      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
-      (except as stated in this section) patent license to make, have made,
-      use, offer to sell, sell, import, and otherwise transfer the Work,
-      where such license applies only to those patent claims licensable
-      by such Contributor that are necessarily infringed by their
-      Contribution(s) alone or by combination of their Contribution(s)
-      with the Work to which such Contribution(s) was submitted. If You
-      institute patent litigation against any entity (including a
-      cross-claim or counterclaim in a lawsuit) alleging that the Work
-      or a Contribution incorporated within the Work constitutes direct
-      or contributory patent infringement, then any patent licenses
-      granted to You under this License for that Work shall terminate
-      as of the date such litigation is filed.
-
-   4. Redistribution. You may reproduce and distribute copies of the
-      Work or Derivative Works thereof in any medium, with or without
-      modifications, and in Source or Object form, provided that You
-      meet the following conditions:
-
-      (a) You must give any other recipients of the Work or
-          Derivative Works a copy of this License; and
-
-      (b) You must cause any modified files to carry prominent notices
-          stating that You changed the files; and
-
-      (c) You must retain, in the Source form of any Derivative Works
-          that You distribute, all copyright, patent, trademark, and
-          attribution notices from the Source form of the Work,
-          excluding those notices that do not pertain to any part of
-          the Derivative Works; and
-
-      (d) If the Work includes a "NOTICE" text file as part of its
-          distribution, then any Derivative Works that You distribute must
-          include a readable copy of the attribution notices contained
-          within such NOTICE file, excluding those notices that do not
-          pertain to any part of the Derivative Works, in at least one
-          of the following places: within a NOTICE text file distributed
-          as part of the Derivative Works; within the Source form or
-          documentation, if provided along with the Derivative Works; or,
-          within a display generated by the Derivative Works, if and
-          wherever such third-party notices normally appear. The contents
-          of the NOTICE file are for informational purposes only and
-          do not modify the License. You may add Your own attribution
-          notices within Derivative Works that You distribute, alongside
-          or as an addendum to the NOTICE text from the Work, provided
-          that such additional attribution notices cannot be construed
-          as modifying the License.
-
-      You may add Your own copyright statement to Your modifications and
-      may provide additional or different license terms and conditions
-      for use, reproduction, or distribution of Your modifications, or
-      for any such Derivative Works as a whole, provided Your use,
-      reproduction, and distribution of the Work otherwise complies with
-      the conditions stated in this License.
-
-   5. Submission of Contributions. Unless You explicitly state otherwise,
-      any Contribution intentionally submitted for inclusion in the Work
-      by You to the Licensor shall be under the terms and conditions of
-      this License, without any additional terms or conditions.
-      Notwithstanding the above, nothing herein shall supersede or modify
-      the terms of any separate license agreement you may have executed
-      with Licensor regarding such Contributions.
-
-   6. Trademarks. This License does not grant permission to use the trade
-      names, trademarks, service marks, or product names of the Licensor,
-      except as required for reasonable and customary use in describing the
-      origin of the Work and reproducing the content of the NOTICE file.
-
-   7. Disclaimer of Warranty. Unless required by applicable law or
-      agreed to in writing, Licensor provides the Work (and each
-      Contributor provides its Contributions) on an "AS IS" BASIS,
-      WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
-      implied, including, without limitation, any warranties or conditions
-      of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
-      PARTICULAR PURPOSE. You are solely responsible for determining the
-      appropriateness of using or redistributing the Work and assume any
-      risks associated with Your exercise of permissions under this License.
-
-   8. Limitation of Liability. In no event and under no legal theory,
-      whether in tort (including negligence), contract, or otherwise,
-      unless required by applicable law (such as deliberate and grossly
-      negligent acts) or agreed to in writing, shall any Contributor be
-      liable to You for damages, including any direct, indirect, special,
-      incidental, or consequential damages of any character arising as a
-      result of this License or out of the use or inability to use the
-      Work (including but not limited to damages for loss of goodwill,
-      work stoppage, computer failure or malfunction, or any and all
-      other commercial damages or losses), even if such Contributor
-      has been advised of the possibility of such damages.
-
-   9. Accepting Warranty or Additional Liability. While redistributing
-      the Work or Derivative Works thereof, You may choose to offer,
-      and charge a fee for, acceptance of support, warranty, indemnity,
-      or other liability obligations and/or rights consistent with this
-      License. However, in accepting such obligations, You may act only
-      on Your own behalf and on Your sole responsibility, not on behalf
-      of any other Contributor, and only if You agree to indemnify,
-      defend, and hold each Contributor harmless for any liability
-      incurred by, or claims asserted against, such Contributor by reason
-      of your accepting any such warranty or additional liability.
-
-   END OF TERMS AND CONDITIONS
-
-   APPENDIX: How to apply the Apache License to your work.
-
-      To apply the Apache License to your work, attach the following
-      boilerplate notice, with the fields enclosed by brackets "[]"
-      replaced with your own identifying information. (Don't include
-      the brackets!)  The text should be enclosed in the appropriate
-      comment syntax for the file format. We also recommend that a
-      file or class name and description of purpose be included on the
-      same "printed page" as the copyright notice for easier
-      identification within third-party archives.
-
-   Copyright [yyyy] [name of copyright owner]
-
-### Non-Open Source Reference Implementations:
-
-#### [MultiPL-E](https://github.com/nuprl/MultiPL-E)
-   - License: BSD 3-Clause License with Machine Learning Restriction
-
-BSD 3-Clause License with Machine Learning Restriction
-
-Copyright (c) 2022, Northeastern University, Oberlin College, Roblox Inc,
-Stevens Institute of Technology, University of Massachusetts Amherst, and
-Wellesley College.
-
-All rights reserved.
-
-Redistribution and use in source and binary forms, with or without
-modification, are permitted provided that the following conditions are met:
-
-1. Redistributions of source code must retain the above copyright notice, this
-   list of conditions and the following disclaimer.
-
-2. Redistributions in binary form must reproduce the above copyright notice,
-   this list of conditions and the following disclaimer in the documentation
-   and/or other materials provided with the distribution.
-
-3. Neither the name of the copyright holder nor the names of its
-   contributors may be used to endorse or promote products derived from
-   this software without specific prior written permission.
-
-4.  The contents of this repository may not be used as training data for any
-    machine learning model, including but not limited to neural networks.
-
-THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS"
-AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE
-IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
-DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE
-FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
-DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
-SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
-CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
-OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
-OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
--- a/Development.md
+++ b/Development.md
@@ -1,331 +1,98 @@
 # Development Guide
-
-This guide is for people working on OpenHands and editing the source code.
-If you wish to contribute your changes, check out the
-[CONTRIBUTING.md](https://github.com/OpenHands/OpenHands/blob/main/CONTRIBUTING.md)
-on how to clone and setup the project initially before moving on. Otherwise,
-you can clone the OpenHands project directly.
-
-## Choose Your Setup
-
-Select your operating system to see the specific setup instructions:
-
- [macOS](#macos-setup)
- [Linux](#linux-setup)
- [Windows WSL](#windows-wsl-setup)
- [Dev Container](#dev-container)
- [Developing in Docker](#developing-in-docker)
- [No sudo access?](#develop-without-sudo-access)
-
---
-
-## macOS Setup
-
-### 1. Install Prerequisites
-
-You'll need the following installed:
-
- **Python 3.12** — `brew install python@3.12` (see the [official Homebrew Python docs](https://docs.brew.sh/Homebrew-and-Python) for details). Make sure `python3.12` is available in your PATH (the `make build` step will verify this).
- **Node.js >= 22** — `brew install node`
- **Poetry >= 1.8** — `brew install poetry`
- **Docker Desktop** — `brew install --cask docker`
-  - After installing, open Docker Desktop → **Settings → Advanced** → Enable **"Allow the default Docker socket to be used"**
-
-### 2. Build and Setup the Environment
-
-```bash
-make build
-```
-
-### 3. Configure the Language Model
-
-OpenHands supports a diverse array of Language Models (LMs) through the powerful [litellm](https://docs.litellm.ai) library.
-
-For the V1 web app, start OpenHands and configure your model and API key in the Settings UI.
-
-If you are running headless or CLI workflows, you can prepare local defaults with:
-
-```bash
-make setup-config
-```
-
-**Note on Alternative Models:**
-See [our documentation](https://docs.openhands.dev/usage/llms) for recommended models.
-
-### 4. Run the Application
-
-```bash
-# Run both backend and frontend
-make run
-
-# Or run separately:
-make start-backend  # Backend only on port 3000
-make start-frontend # Frontend only on port 3001
-```
-
-These targets serve the current OpenHands V1 API by default. In the codebase, `make start-backend` runs `openhands.server.listen:app`, and that app includes the `openhands/app_server` V1 routes unless `ENABLE_V1=0`.
-
---
-
-## Linux Setup
-
-This guide covers Ubuntu/Debian. For other distributions, adapt the package manager commands accordingly.
-
-### 1. Install Prerequisites
-
-```bash
-# Update package list
-sudo apt update
-
-# Install system dependencies
-sudo apt install -y build-essential curl netcat software-properties-common
-
-# Install Python 3.12
-# Ubuntu 24.04+ and Debian 13+ ship with Python 3.12 — skip the PPA step if
-# python3.12 --version already works on your system.
-# The deadsnakes PPA is Ubuntu-only and needed for Ubuntu 22.04 or older:
-sudo add-apt-repository -y ppa:deadsnakes/ppa
-sudo apt update
-sudo apt install -y python3.12 python3.12-dev python3.12-venv
-
-# Install Node.js 22.x
-curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash -
-sudo apt install -y nodejs
-
-# Install Poetry
-curl -sSL https://install.python-poetry.org | python3 -
-
-# Add Poetry to your PATH
-echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc
-source ~/.bashrc
-
-# Install Docker
-# Follow the official guide: https://docs.docker.com/engine/install/ubuntu/
-# Quick version:
-sudo install -m 0755 -d /etc/apt/keyrings
-curl -fsSL https://download.docker.com/linux/ubuntu/gpg -o /etc/apt/keyrings/docker.asc
-sudo chmod a+r /etc/apt/keyrings/docker.asc
-echo "deb [arch=$(dpkg --print-architecture) signed-by=/etc/apt/keyrings/docker.asc] https://download.docker.com/linux/ubuntu $(. /etc/os-release && echo "$VERSION_CODENAME") stable" | sudo tee /etc/apt/sources.list.d/docker.list > /dev/null
-sudo apt update
-sudo apt install -y docker-ce docker-ce-cli containerd.io docker-buildx-plugin docker-compose-plugin
-sudo usermod -aG docker $USER
-# Log out and back in for Docker group changes to take effect
-```
-
-### 2. Build and Setup the Environment
-
-```bash
-make build
-```
-
-### 3. Configure the Language Model
-
-See the [macOS section above](#3-configure-the-language-model) for guidance: configure your model and API key in the Settings UI.
-
-### 4. Run the Application
-
-```bash
-# Run both backend and frontend
-make run
-
-# Or run separately:
-make start-backend  # Backend only on port 3000
-make start-frontend # Frontend only on port 3001
-```
-
---
-
-## Windows WSL Setup
-
-WSL2 with Ubuntu is recommended. The setup is similar to Linux, with a few WSL-specific considerations.
-
-### 1. Install WSL2
-
-**Option A: Windows 11 (Microsoft Store)**
-The easiest way on Windows 11:
-1. Open the **Microsoft Store** app
-2. Search for **"Ubuntu 22.04 LTS"** or **"Ubuntu"**
-3. Click **Install**
-4. Launch Ubuntu from the Start menu
-
-**Option B: PowerShell**
-```powershell
-# Run this in PowerShell as Administrator
-wsl --install -d Ubuntu-22.04
-```
-
-After installation, restart your computer and open Ubuntu.
-
-### 2. Install Prerequisites (in WSL Ubuntu)
-
-Follow [Step 1 from the Linux setup](#1-install-prerequisites-1) to install system dependencies, Python 3.12, Node.js, and Poetry. Skip the Docker installation — Docker is provided through Docker Desktop below.
-
-### 3. Configure Docker for WSL2
-
-1. Install [Docker Desktop for Windows](https://www.docker.com/products/docker-desktop)
-2. Open Docker Desktop > Settings > General
-3. Enable: "Use the WSL 2 based engine"
-4. Go to Settings > Resources > WSL Integration
-5. Enable integration with your Ubuntu distribution
-
-**Important:** Keep your project files in the WSL filesystem (e.g., `~/workspace/openhands`), not in `/mnt/c`. Files accessed via `/mnt/c` will be significantly slower.
-
-### 4. Build and Setup the Environment
-
-```bash
-make build
-```
-
-### 5. Configure the Language Model
-
-See the [macOS section above](#3-configure-the-language-model) for the current V1 guidance: configure your model and API key in the Settings UI for the web app, and use `make setup-config` only for headless or CLI workflows.
-
-### 6. Run the Application
-
-```bash
-# Run both backend and frontend
-make run
-
-# Or run separately:
-make start-backend  # Backend only on port 3000
-make start-frontend # Frontend only on port 3001
-```
-
-Access the frontend at `http://localhost:3001` from your Windows browser.
-
---
-
-## Dev Container
-
-There is a [dev container](https://containers.dev/) available which provides a
-pre-configured environment with all the necessary dependencies installed if you
-are using a [supported editor or tool](https://containers.dev/supporting). For
-example, if you are using Visual Studio Code (VS Code) with the
-[Dev Containers](https://marketplace.visualstudio.com/items?itemName=ms-vscode-remote.remote-containers)
-extension installed, you can open the project in a dev container by using the
-_Dev Container: Reopen in Container_ command from the Command Palette
-(Ctrl+Shift+P).
-
---
-
-## Developing in Docker
-
-If you don't want to install dependencies on your host machine, you can develop inside a Docker container.
-
-### Quick Start
-
-```bash
-make docker-dev
-```
-
-For more details, see the [dev container documentation](./containers/dev/README.md).
-
-### Alternative: Docker Run
-
-If you just want to run OpenHands without setting up a dev environment:
-
-```bash
-make docker-run
-```
-
-If you don't have `make` installed, run:
-
-```bash
-cd ./containers/dev
-./dev.sh
-```
-
---
-
-## Develop without sudo access
-
-If you want to develop without system admin/sudo access to upgrade/install `Python` and/or `NodeJS`, you can use
-`conda` or `mamba` to manage the packages for you:
+This guide is for people working on OpenDevin and editing the source code.
+If you wish to contribute your changes, check out the [CONTRIBUTING.md](https://github.com/OpenDevin/OpenDevin/blob/main/CONTRIBUTING.md) on how to clone and setup the project initially before moving on.
+Otherwise, you can clone the OpenDevin project directly.
+
+## Start the server for development
+### 1. Requirements
+* Linux, Mac OS, or [WSL on Windows](https://learn.microsoft.com/en-us/windows/wsl/install)  [ Ubuntu <= 22.04]
+* [Docker](https://docs.docker.com/engine/install/) (For those on MacOS, make sure to allow the default Docker socket to be used from advanced settings!)
+* [Python](https://www.python.org/downloads/) = 3.11
+* [NodeJS](https://nodejs.org/en/download/package-manager) >= 18.17.1
+* [Poetry](https://python-poetry.org/docs/#installing-with-the-official-installer) >= 1.8
+* netcat => sudo apt-get install netcat
+
+Make sure you have all these dependencies installed before moving on to `make build`.
+
+#### Develop without sudo access
+If you want to develop without system admin/sudo access to upgrade/install `Python` and/or `NodeJs`, you can use `conda` or `mamba` to manage the packages for you:

 ```bash
 # Download and install Mamba (a faster version of conda)
 curl -L -O "https://github.com/conda-forge/miniforge/releases/latest/download/Miniforge3-$(uname)-$(uname -m).sh"
 bash Miniforge3-$(uname)-$(uname -m).sh

-# Install Python 3.12, nodejs, and poetry
-mamba install python=3.12
+# Install Python 3.11, nodejs, and poetry
+mamba install python=3.11
 mamba install conda-forge::nodejs
 mamba install conda-forge::poetry
 ```

---
-
-## Running OpenHands with OpenHands
-
-You can use OpenHands to develop and improve OpenHands itself!
-
-### Quick Start
+### 2. Build and Setup The Environment
+Begin by building the project which includes setting up the environment and installing dependencies. This step ensures that OpenDevin is ready to run on your system:

 ```bash
-export INSTALL_DOCKER=0
-export RUNTIME=local
-make build && make run
+make build
 ```

-Access the interface at:
- Local development: http://localhost:3001
- Remote/cloud environments: Use the appropriate external URL
+### 3. Configuring the Language Model
+OpenDevin supports a diverse array of Language Models (LMs) through the powerful [litellm](https://docs.litellm.ai) library. By default, we've chosen the mighty GPT-4 from OpenAI as our go-to model, but the world is your oyster! You can unleash the potential of Anthropic's suave Claude, the enigmatic Llama, or any other LM that piques your interest.

-For external access:
+To configure the LM of your choice, run:
+       
+   ```bash
+   make setup-config
+   ```
+   
+   This command will prompt you to enter the LLM API key, model name, and other variables ensuring that OpenDevin is tailored to your specific needs. Note that the model name will apply only when you run headless. If you use the UI, please set the model in the UI.
+   
+   Note: If you have previously run OpenDevin using the docker command, you may have already set some environmental variables in your terminal. The final configurations are set from highest to lowest priority:
+   Environment variables > config.toml variables > default variables
+
+**Note on Alternative Models:**
+Some alternative models may prove more challenging to tame than others. Fear not, brave adventurer! We shall soon unveil LLM-specific documentation to guide you on your quest. 
+And if you've already mastered the art of wielding a model other than OpenAI's GPT, we encourage you to share your setup instructions with us by creating instructions and adding it [to our documentation](https://github.com/OpenDevin/OpenDevin/tree/main/docs/modules/usage/llms).
+
+For a full list of the LM providers and models available, please consult the [litellm documentation](https://docs.litellm.ai/docs/providers).
+
+### 4. Running the application
+#### Option A: Run the Full Application
+Once the setup is complete, launching OpenDevin is as simple as running a single command. This command starts both the backend and frontend servers seamlessly, allowing you to interact with OpenDevin:
 ```bash
-make run FRONTEND_PORT=12000 FRONTEND_HOST=0.0.0.0 BACKEND_HOST=0.0.0.0
+make run
 ```

---
+#### Option B: Individual Server Startup
+- **Start the Backend Server:** If you prefer, you can start the backend server independently to focus on backend-related tasks or configurations.
+    ```bash
+    make start-backend
+    ```

-## LLM Debugging
+- **Start the Frontend Server:** Similarly, you can start the frontend server on its own to work on frontend-related components or interface enhancements.
+    ```bash
+    make start-frontend
+    ```

-If you encounter issues with the Language Model, enable debug logging:
-
-```bash
-export DEBUG=1
-# Restart the backend
-make start-backend
-```
-
-Logs will be saved to `logs/llm/CURRENT_DATE/` for troubleshooting.
-
---
-
-## Testing
-
-### Unit Tests
-
-```bash
-poetry run pytest ./tests/unit/test_*.py
-```
-
---
-
-## Adding Dependencies
-
-1. Add your dependency in `pyproject.toml` or use `poetry add xxx`
-2. Update the lock file: `poetry lock --no-update`
-
---
-
-## Help
+### 6. LLM Debugging
+If you encounter any issues with the Language Model (LM) or you're simply curious, you can inspect the actual LLM prompts and responses. To do so, export DEBUG=1 in the environment and restart the backend.
+OpenDevin will then log the prompts and responses in the logs/llm/CURRENT_DATE directory, allowing you to identify the causes.

+### 7. Help
+Need assistance or information on available targets and commands? The help command provides all the necessary guidance to ensure a smooth experience with OpenDevin.
 ```bash
 make help
+ ```
+
+### 8. Testing
+#### Unit tests
+
+```bash
+poetry run pytest ./tests/unit/test_sandbox.py
 ```

---
+#### Integration tests
+Please refer to [this README](./tests/integration/README.md) for details.

-## Key Documentation Resources
-
- [/README.md](./README.md): Main project overview, features, and basic setup instructions
- [/Development.md](./Development.md) (this file): Comprehensive guide for developers working on OpenHands
- [/CONTRIBUTING.md](./CONTRIBUTING.md): Guidelines for contributing to the project, including code style and PR process
- [DOC_STYLE_GUIDE.md](https://github.com/OpenHands/docs/blob/main/openhands/DOC_STYLE_GUIDE.md): Standards for writing and maintaining project documentation
- [/openhands/app_server/README.md](./openhands/app_server/README.md): Current V1 application server implementation and REST API modules
- [/frontend/README.md](./frontend/README.md): Frontend React application setup and development guide
- [/containers/README.md](./containers/README.md): Information about Docker containers and deployment
- [/tests/unit/README.md](./tests/unit/README.md): Guide to writing and running unit tests
- [OpenHands/benchmarks](https://github.com/OpenHands/benchmarks): Documentation for the evaluation framework and benchmarks
- [/skills/README.md](./skills/README.md): Information about the skills architecture and implementation
+### 9. Add or update dependency
+1. Add your dependency in `pyproject.toml` or use `poetry add xxx`
+2. Update the poetry.lock file via `poetry lock --no-update`
--- a/ISSUE_TRIAGE.md
+++ b/ISSUE_TRIAGE.md
@@ -1,27 +0,0 @@
-# Issue Triage
-These are the procedures and guidelines on how issues are triaged in this repo by the maintainers.
-
-## General
-* All issues must be tagged with **enhancement**, **bug** or **troubleshooting/help**.
-* Issues may be tagged with what it relates to (**llm**, **app tab**, **UI/UX**, etc.).
-
-## Severity
-* **High**: High visibility issues or affecting many users.
-* **Critical**: Affecting all users or potential security issues.
-
-## Difficulty
-* Issues good for newcomers may be tagged with **good first issue**.
-
-## Not Enough Information
-* User is asked to provide more information (logs, how to reproduce, etc.) when the issue is not clear.
-* If an issue is unclear and the author does not provide more information or respond to a request,
-the issue may be closed as **not planned** (Usually after a week).
-
-## Multiple Requests/Fixes in One Issue
-* These issues will be narrowed down to one request/fix so the issue is more easily tracked and fixed.
-* Issues may be broken down into multiple issues if required.
-
-## Stale and Auto Closures
-* In order to keep a maintainable backlog, issues that have no activity within 40 days are automatically marked as **Stale**.
-* If issues marked as **Stale** continue to have no activity for 10 more days, they will automatically be closed as not planned.
-* Issues may be reopened by maintainers if deemed important.
--- a/9
+++ b/9
@@ -1,12 +1,7 @@
-Portions of this software are licensed as follows:
-* All content that resides under the enterprise/ directory is licensed under the license defined in "enterprise/LICENSE".
-* Content outside of the above mentioned directories or restrictions above is available under the MIT license as defined below.
-
+The MIT License (MIT)
 =====================

-The MIT License (MIT)
-
-Copyright © 2025
+Copyright © 2023

 Permission is hereby granted, free of charge, to any person
 obtaining a copy of this software and associated documentation
--- a/MANIFEST.in
+++ b/MANIFEST.in
@@ -1,5 +0,0 @@
-# Exclude all Python bytecode files
-global-exclude *.pyc
-
-# Exclude Python cache directories
-global-exclude __pycache__
--- a/256
+++ b/256
@@ -1,26 +1,16 @@
-SHELL=/usr/bin/env bash
-# Makefile for OpenHands project
+SHELL=/bin/bash
+# Makefile for OpenDevin project

 # Variables
-BACKEND_HOST ?= "127.0.0.1"
-BACKEND_PORT ?= 3000
-BACKEND_HOST_PORT = "$(BACKEND_HOST):$(BACKEND_PORT)"
-FRONTEND_HOST ?= "127.0.0.1"
-FRONTEND_PORT ?= 3001
+DOCKER_IMAGE = ghcr.io/opendevin/sandbox:main
+BACKEND_PORT = 3000
+BACKEND_HOST = "127.0.0.1:$(BACKEND_PORT)"
+FRONTEND_PORT = 3001
 DEFAULT_WORKSPACE_DIR = "./workspace"
 DEFAULT_MODEL = "gpt-4o"
 CONFIG_FILE = config.toml
 PRE_COMMIT_CONFIG_PATH = "./dev_config/python/.pre-commit-config.yaml"
-PYTHON_MIN_VERSION = 3.12
-PYTHON_MAX_VERSION = 3.14
-PYTHON_CANDIDATES ?= python3.13 python3.12 python3
-PYTHON ?= $(shell for cmd in $(PYTHON_CANDIDATES); do \
-	if command -v $$cmd > /dev/null 2>&1 && $$cmd -c 'import sys; raise SystemExit(0 if ((3, 12) <= sys.version_info[:2] < (3, 14)) else 1)' > /dev/null 2>&1; then \
-		echo $$cmd; \
-		exit 0; \
-	fi; \
- done)
-KIND_CLUSTER_NAME = "local-hands"
+PYTHON_VERSION = 3.11

 # ANSI color codes
 GREEN=$(shell tput -Txterm setaf 2)
@@ -33,6 +23,9 @@ RESET=$(shell tput -Txterm sgr0)
 build:
 	@echo "$(GREEN)Building project...$(RESET)"
 	@$(MAKE) -s check-dependencies
+ifeq ($(INSTALL_DOCKER),)
+	@$(MAKE) -s pull-docker-image
+endif
 	@$(MAKE) -s install-python-dependencies
 	@$(MAKE) -s install-frontend-dependencies
 	@$(MAKE) -s install-pre-commit-hooks
@@ -49,7 +42,6 @@ ifeq ($(INSTALL_DOCKER),)
 	@$(MAKE) -s check-docker
 endif
 	@$(MAKE) -s check-poetry
-	@$(MAKE) -s check-tmux
 	@echo "$(GREEN)Dependencies checked successfully.$(RESET)"

 check-system:
@@ -71,10 +63,10 @@ check-system:

 check-python:
 	@echo "$(YELLOW)Checking Python installation...$(RESET)"
-	@if [ -n "$(PYTHON)" ]; then \
-		echo "$(BLUE)$$($(PYTHON) --version) is already installed (using $(PYTHON)).$(RESET)"; \
+	@if command -v python$(PYTHON_VERSION) > /dev/null; then \
+		echo "$(BLUE)$(shell python$(PYTHON_VERSION) --version) is already installed.$(RESET)"; \
 	else \
-		echo "$(RED)A compatible Python interpreter (>= $(PYTHON_MIN_VERSION), < $(PYTHON_MAX_VERSION)) is required. Please install Python 3.12 or 3.13 to continue.$(RESET)"; \
+		echo "$(RED)Python $(PYTHON_VERSION) is not installed. Please install Python $(PYTHON_VERSION) to continue.$(RESET)"; \
 		exit 1; \
 	fi

@@ -92,10 +84,10 @@ check-nodejs:
 	@if command -v node > /dev/null; then \
 		NODE_VERSION=$(shell node --version | sed -E 's/v//g'); \
 		IFS='.' read -r -a NODE_VERSION_ARRAY <<< "$$NODE_VERSION"; \
-		if [ "$${NODE_VERSION_ARRAY[0]}" -ge 22 ]; then \
+		if [ "$${NODE_VERSION_ARRAY[0]}" -gt 18 ] || ([ "$${NODE_VERSION_ARRAY[0]}" -eq 18 ] && [ "$${NODE_VERSION_ARRAY[1]}" -gt 17 ]) || ([ "$${NODE_VERSION_ARRAY[0]}" -eq 18 ] && [ "$${NODE_VERSION_ARRAY[1]}" -eq 17 ] && [ "$${NODE_VERSION_ARRAY[2]}" -ge 1 ]); then \
 			echo "$(BLUE)Node.js $$NODE_VERSION is already installed.$(RESET)"; \
 		else \
-			echo "$(RED)Node.js 22.x or later is required. Please install Node.js 22.x or later to continue.$(RESET)"; \
+			echo "$(RED)Node.js 18.17.1 or later is required. Please install Node.js 18.17.1 or later to continue.$(RESET)"; \
 			exit 1; \
 		fi; \
 	else \
@@ -112,98 +104,82 @@ check-docker:
 		exit 1; \
 	fi

-check-tmux:
-	@echo "$(YELLOW)Checking tmux installation...$(RESET)"
-	@if command -v tmux > /dev/null; then \
-		echo "$(BLUE)$(shell tmux -V) is already installed.$(RESET)"; \
-	else \
-		echo "$(YELLOW)╔════════════════════════════════════════════════════════════════════════════╗$(RESET)"; \
-		echo "$(YELLOW)║ OPTIONAL: tmux is not installed.                                          ║$(RESET)"; \
-		echo "$(YELLOW)║ Some advanced terminal features may not work without tmux.                ║$(RESET)"; \
-		echo "$(YELLOW)║ You can install it if needed, but it's not required for development.      ║$(RESET)"; \
-		echo "$(YELLOW)╚════════════════════════════════════════════════════════════════════════════╝$(RESET)"; \
-	fi
-
 check-poetry:
 	@echo "$(YELLOW)Checking Poetry installation...$(RESET)"
-	@if [ -z "$(PYTHON)" ]; then \
-		echo "$(RED)A compatible Python interpreter (>= $(PYTHON_MIN_VERSION), < $(PYTHON_MAX_VERSION)) is required. Please install Python 3.12 or 3.13 to continue.$(RESET)"; \
-		exit 1; \
-	elif command -v poetry > /dev/null; then \
+	@if command -v poetry > /dev/null; then \
 		POETRY_VERSION=$(shell poetry --version 2>&1 | sed -E 's/Poetry \(version ([0-9]+\.[0-9]+\.[0-9]+)\)/\1/'); \
 		IFS='.' read -r -a POETRY_VERSION_ARRAY <<< "$$POETRY_VERSION"; \
-		if [ $${POETRY_VERSION_ARRAY[0]} -gt 1 ] || ([ $${POETRY_VERSION_ARRAY[0]} -eq 1 ] && [ $${POETRY_VERSION_ARRAY[1]} -ge 8 ]); then \
+		if [ $${POETRY_VERSION_ARRAY[0]} -ge 1 ] && [ $${POETRY_VERSION_ARRAY[1]} -ge 8 ]; then \
 			echo "$(BLUE)$(shell poetry --version) is already installed.$(RESET)"; \
 		else \
 			echo "$(RED)Poetry 1.8 or later is required. You can install poetry by running the following command, then adding Poetry to your PATH:"; \
-			echo "$(RED) curl -sSL https://install.python-poetry.org | $(PYTHON) -$(RESET)"; \
+			echo "$(RED) curl -sSL https://install.python-poetry.org | python$(PYTHON_VERSION) -$(RESET)"; \
 			echo "$(RED)More detail here: https://python-poetry.org/docs/#installing-with-the-official-installer$(RESET)"; \
 			exit 1; \
 		fi; \
 	else \
 		echo "$(RED)Poetry is not installed. You can install poetry by running the following command, then adding Poetry to your PATH:"; \
-		echo "$(RED) curl -sSL https://install.python-poetry.org | $(PYTHON) -$(RESET)"; \
+		echo "$(RED) curl -sSL https://install.python-poetry.org | python$(PYTHON_VERSION) -$(RESET)"; \
 		echo "$(RED)More detail here: https://python-poetry.org/docs/#installing-with-the-official-installer$(RESET)"; \
 		exit 1; \
 	fi

-install-python-dependencies: check-python
+pull-docker-image:
+	@echo "$(YELLOW)Pulling Docker image...$(RESET)"
+	@docker pull $(DOCKER_IMAGE)
+	@echo "$(GREEN)Docker image pulled successfully.$(RESET)"
+
+install-python-dependencies:
 	@echo "$(GREEN)Installing Python dependencies...$(RESET)"
 	@if [ -z "${TZ}" ]; then \
 		echo "Defaulting TZ (timezone) to UTC"; \
 		export TZ="UTC"; \
 	fi
-	poetry env use $(PYTHON)
+	poetry env use python$(PYTHON_VERSION)
 	@if [ "$(shell uname)" = "Darwin" ]; then \
 		echo "$(BLUE)Installing chroma-hnswlib...$(RESET)"; \
 		export HNSWLIB_NO_NATIVE=1; \
 		poetry run pip install chroma-hnswlib; \
 	fi
-	@if [ -n "${POETRY_GROUP}" ]; then \
-		echo "Installing only POETRY_GROUP=${POETRY_GROUP}"; \
-		poetry install --only $${POETRY_GROUP}; \
+	@poetry install
+	@if [ -f "/etc/manjaro-release" ]; then \
+		echo "$(BLUE)Detected Manjaro Linux. Installing Playwright dependencies...$(RESET)"; \
+		poetry run pip install playwright; \
+		poetry run playwright install chromium; \
 	else \
-		poetry install --with dev,test,runtime; \
-	fi
-	@if [ "${INSTALL_PLAYWRIGHT}" != "false" ] && [ "${INSTALL_PLAYWRIGHT}" != "0" ]; then \
-		if [ -f "/etc/manjaro-release" ]; then \
-			echo "$(BLUE)Detected Manjaro Linux. Installing Playwright dependencies...$(RESET)"; \
-			poetry run pip install playwright; \
-			poetry run playwright install chromium; \
+		if [ ! -f cache/playwright_chromium_is_installed.txt ]; then \
+			echo "Running playwright install --with-deps chromium..."; \
+			poetry run playwright install --with-deps chromium; \
+			mkdir -p cache; \
+			touch cache/playwright_chromium_is_installed.txt; \
 		else \
-			if [ ! -f cache/playwright_chromium_is_installed.txt ]; then \
-				echo "Running playwright install --with-deps chromium..."; \
-				poetry run playwright install --with-deps chromium; \
-				mkdir -p cache; \
-				touch cache/playwright_chromium_is_installed.txt; \
-			else \
-				echo "Setup already done. Skipping playwright installation."; \
-			fi \
+			echo "Setup already done. Skipping playwright installation."; \
 		fi \
-	else \
-		echo "Skipping Playwright installation (INSTALL_PLAYWRIGHT=${INSTALL_PLAYWRIGHT})."; \
 	fi
 	@echo "$(GREEN)Python dependencies installed successfully.$(RESET)"

-install-frontend-dependencies: check-npm check-nodejs
+install-frontend-dependencies:
 	@echo "$(YELLOW)Setting up frontend environment...$(RESET)"
 	@echo "$(YELLOW)Detect Node.js version...$(RESET)"
 	@cd frontend && node ./scripts/detect-node-version.js
-	echo "$(BLUE)Installing frontend dependencies with npm...$(RESET)"
-	@cd frontend && npm install
+	@cd frontend && \
+		echo "$(BLUE)Installing frontend dependencies with npm...$(RESET)" && \
+		npm install && \
+		echo "$(BLUE)Running make-i18n with npm...$(RESET)" && \
+		npm run make-i18n
 	@echo "$(GREEN)Frontend dependencies installed successfully.$(RESET)"

-install-pre-commit-hooks: check-python check-poetry install-python-dependencies
+install-pre-commit-hooks:
 	@echo "$(YELLOW)Installing pre-commit hooks...$(RESET)"
 	@git config --unset-all core.hooksPath || true
 	@poetry run pre-commit install --config $(PRE_COMMIT_CONFIG_PATH)
 	@echo "$(GREEN)Pre-commit hooks installed successfully.$(RESET)"

-lint-backend: install-pre-commit-hooks
+lint-backend:
 	@echo "$(YELLOW)Running linters...$(RESET)"
-	@poetry run pre-commit run --all-files --show-diff-on-failure --config $(PRE_COMMIT_CONFIG_PATH)
+	@poetry run pre-commit run --files opendevin/**/* agenthub/**/* evaluation/**/* --show-diff-on-failure --config $(PRE_COMMIT_CONFIG_PATH)

-lint-frontend: install-frontend-dependencies
+lint-frontend:
 	@echo "$(YELLOW)Running linters for frontend...$(RESET)"
 	@cd frontend && npm run lint

@@ -211,40 +187,6 @@ lint:
 	@$(MAKE) -s lint-frontend
 	@$(MAKE) -s lint-backend

-kind:
-	@echo "$(YELLOW)Checking if kind is installed...$(RESET)"
-	@if ! command -v kind > /dev/null; then \
-		echo "$(RED)kind is not installed. Please install kind with `brew install kind` to continue$(RESET)"; \
-		exit 1; \
-	else \
-		echo "$(BLUE)kind $(shell kind version) is already installed.$(RESET)"; \
-	fi
-	@echo "$(YELLOW)Checking if kind cluster '$(KIND_CLUSTER_NAME)' already exists...$(RESET)"
-	@if kind get clusters | grep -q "^$(KIND_CLUSTER_NAME)$$"; then \
-		echo "$(BLUE)Kind cluster '$(KIND_CLUSTER_NAME)' already exists.$(RESET)"; \
-		kubectl config use-context kind-$(KIND_CLUSTER_NAME); \
-	else \
-		echo "$(YELLOW)Creating kind cluster '$(KIND_CLUSTER_NAME)'...$(RESET)"; \
-		kind create cluster --name $(KIND_CLUSTER_NAME) --config kind/cluster.yaml; \
-	fi
-	@echo "$(YELLOW)Checking if mirrord is installed...$(RESET)"
-	@if ! command -v mirrord > /dev/null; then \
-		echo "$(RED)mirrord is not installed. Please install mirrord with `brew install metalbear-co/mirrord/mirrord` to continue$(RESET)"; \
-		exit 1; \
-	else \
-		echo "$(BLUE)mirrord $(shell mirrord --version) is already installed.$(RESET)"; \
-	fi
-	@echo "$(YELLOW)Installing k8s mirrord resources...$(RESET)"
-	@kubectl apply -f kind/manifests
-	@echo "$(GREEN)Mirrord resources installed successfully.$(RESET)"
-	@echo "$(YELLOW)Waiting for Mirrord pod to be ready.$(RESET)"
-	@sleep 5
-	@kubectl wait --for=condition=Available deployment/ubuntu-dev
-	@echo "$(YELLOW)Waiting for Nginx to be ready.$(RESET)"
-	@kubectl -n ingress-nginx wait --for=condition=Available deployment/ingress-nginx-controller
-	@echo "$(YELLOW)Running make run inside of mirrord.$(RESET)"
-	@mirrord exec --target deployment/ubuntu-dev -- make run
-
 test-frontend:
 	@echo "$(YELLOW)Running tests for frontend...$(RESET)"
 	@cd frontend && npm run test
@@ -254,24 +196,17 @@ test:

 build-frontend:
 	@echo "$(YELLOW)Building frontend...$(RESET)"
-	@cd frontend && npm run prepare && npm run build
+	@cd frontend && npm run build

 # Start backend
 start-backend:
 	@echo "$(YELLOW)Starting backend...$(RESET)"
-	@poetry run uvicorn openhands.server.listen:app --host $(BACKEND_HOST) --port $(BACKEND_PORT) --reload --reload-exclude "./workspace"
+	@poetry run uvicorn opendevin.server.listen:app --port $(BACKEND_PORT) --reload --reload-exclude "workspace/*"

 # Start frontend
 start-frontend:
 	@echo "$(YELLOW)Starting frontend...$(RESET)"
-	@cd frontend && \
-	if grep -qi microsoft /proc/version 2>/dev/null; then \
-		echo "Detected WSL environment. Using 'dev_wsl'"; \
-		SCRIPT=dev_wsl; \
-	else \
-		SCRIPT=dev; \
-	fi; \
-	VITE_BACKEND_HOST=$(BACKEND_HOST_PORT) VITE_FRONTEND_PORT=$(FRONTEND_PORT) npm run $$SCRIPT -- --port $(FRONTEND_PORT) --host $(BACKEND_HOST)
+	@cd frontend && VITE_BACKEND_HOST=$(BACKEND_HOST) VITE_FRONTEND_PORT=$(FRONTEND_PORT) npm run start

 # Common setup for running the app (non-callable)
 _run_setup:
@@ -281,7 +216,7 @@ _run_setup:
 	fi
 	@mkdir -p logs
 	@echo "$(YELLOW)Starting backend server...$(RESET)"
-	@poetry run uvicorn openhands.server.listen:app --host $(BACKEND_HOST) --port $(BACKEND_PORT) &
+	@poetry run uvicorn opendevin.server.listen:app --port $(BACKEND_PORT) &
 	@echo "$(YELLOW)Waiting for the backend to start...$(RESET)"
 	@until nc -z localhost $(BACKEND_PORT); do sleep 0.1; done
 	@echo "$(GREEN)Backend started successfully.$(RESET)"
@@ -290,23 +225,15 @@ _run_setup:
 run:
 	@echo "$(YELLOW)Running the app...$(RESET)"
 	@$(MAKE) -s _run_setup
-	@$(MAKE) -s start-frontend
+	@cd frontend && echo "$(BLUE)Starting frontend with npm...$(RESET)" && npm run start -- --port $(FRONTEND_PORT)
 	@echo "$(GREEN)Application started successfully.$(RESET)"

-# Run the app (in docker)
-docker-run: WORKSPACE_BASE ?= $(PWD)/workspace
-docker-run:
-	@if [ -f /.dockerenv ]; then \
-		echo "Running inside a Docker container. Exiting..."; \
-		exit 0; \
-	else \
-		echo "$(YELLOW)Running the app in Docker $(OPTIONS)...$(RESET)"; \
-		export WORKSPACE_BASE=${WORKSPACE_BASE}; \
-		export SANDBOX_USER_ID=$(shell id -u); \
-		export DATE=$(shell date +%Y%m%d%H%M%S); \
-		docker compose up $(OPTIONS); \
-	fi
-
+# Run the app (WSL mode)
+run-wsl:
+	@echo "$(YELLOW)Running the app in WSL mode...$(RESET)"
+	@$(MAKE) -s _run_setup
+	@cd frontend && echo "$(BLUE)Starting frontend with npm (WSL mode)...$(RESET)" && npm run dev_wsl -- --port $(FRONTEND_PORT)
+	@echo "$(GREEN)Application started successfully in WSL mode.$(RESET)"

 # Setup config.toml
 setup-config:
@@ -322,6 +249,16 @@ setup-config-prompts:
 	 workspace_dir=$${workspace_dir:-$(DEFAULT_WORKSPACE_DIR)}; \
 	 echo "workspace_base=\"$$workspace_dir\"" >> $(CONFIG_FILE).tmp

+	@read -p "Do you want to persist the sandbox container? [true/false] [default: false]: " persist_sandbox; \
+	 persist_sandbox=$${persist_sandbox:-false}; \
+	 if [ "$$persist_sandbox" = "true" ]; then \
+		 read -p "Enter a password for the sandbox container: " ssh_password; \
+		 echo "ssh_password=\"$$ssh_password\"" >> $(CONFIG_FILE).tmp; \
+		 echo "persist_sandbox=$$persist_sandbox" >> $(CONFIG_FILE).tmp; \
+	 else \
+		echo "persist_sandbox=$$persist_sandbox" >> $(CONFIG_FILE).tmp; \
+	 fi
+
 	@echo "" >> $(CONFIG_FILE).tmp

 	@echo "[llm]" >> $(CONFIG_FILE).tmp
@@ -335,30 +272,36 @@ setup-config-prompts:
 	@read -p "Enter your LLM base URL [mostly used for local LLMs, leave blank if not needed - example: http://localhost:5001/v1/]: " llm_base_url; \
 	 if [[ ! -z "$$llm_base_url" ]]; then echo "base_url=\"$$llm_base_url\"" >> $(CONFIG_FILE).tmp; fi

-setup-config-basic:
-	@printf '%s\n' \
-	'[core]' \
-	'workspace_base="./workspace"' \
-	> config.toml
-	@echo "$(GREEN)config.toml created.$(RESET)"
+	@echo "Enter your LLM Embedding Model"; \
+		echo "Choices are:"; \
+		echo "  - openai"; \
+		echo "  - azureopenai"; \
+		echo "  - Embeddings available only with OllamaEmbedding:"; \
+		echo "    - llama2"; \
+		echo "    - mxbai-embed-large"; \
+		echo "    - nomic-embed-text"; \
+		echo "    - all-minilm"; \
+		echo "    - stable-code"; \
+		echo "  - Leave blank to default to 'BAAI/bge-small-en-v1.5' via huggingface"; \
+		read -p "> " llm_embedding_model; \
+		echo "embedding_model=\"$$llm_embedding_model\"" >> $(CONFIG_FILE).tmp; \
+		if [ "$$llm_embedding_model" = "llama2" ] || [ "$$llm_embedding_model" = "mxbai-embed-large" ] || [ "$$llm_embedding_model" = "nomic-embed-text" ] || [ "$$llm_embedding_model" = "all-minilm" ] || [ "$$llm_embedding_model" = "stable-code" ]; then \
+			read -p "Enter the local model URL for the embedding model (will set llm.embedding_base_url): " llm_embedding_base_url; \
+				echo "embedding_base_url=\"$$llm_embedding_base_url\"" >> $(CONFIG_FILE).tmp; \
+		elif [ "$$llm_embedding_model" = "azureopenai" ]; then \
+			read -p "Enter the Azure endpoint URL (will overwrite llm.base_url): " llm_base_url; \
+				echo "base_url=\"$$llm_base_url\"" >> $(CONFIG_FILE).tmp; \
+			read -p "Enter the Azure LLM Embedding Deployment Name: " llm_embedding_deployment_name; \
+				echo "embedding_deployment_name=\"$$llm_embedding_deployment_name\"" >> $(CONFIG_FILE).tmp; \
+			read -p "Enter the Azure API Version: " llm_api_version; \
+				echo "api_version=\"$$llm_api_version\"" >> $(CONFIG_FILE).tmp; \
+		fi

-openhands-cloud-run:
-	@$(MAKE) run BACKEND_HOST="0.0.0.0" BACKEND_PORT="12000" FRONTEND_HOST="0.0.0.0" FRONTEND_PORT="12001"
-
-# Develop in container
-docker-dev:
-	@if [ -f /.dockerenv ]; then \
-		echo "Running inside a Docker container. Exiting..."; \
-		exit 0; \
-	else \
-		echo "$(YELLOW)Build and run in Docker $(OPTIONS)...$(RESET)"; \
-		./containers/dev/dev.sh $(OPTIONS); \
-	fi

 # Clean up all caches
 clean:
 	@echo "$(YELLOW)Cleaning up caches...$(RESET)"
-	@rm -rf openhands/.cache
+	@rm -rf opendevin/.cache
 	@echo "$(GREEN)Caches cleaned up successfully.$(RESET)"

 # Help
@@ -367,16 +310,13 @@ help:
 	@echo "Targets:"
 	@echo "  $(GREEN)build$(RESET)               - Build project, including environment setup and dependencies."
 	@echo "  $(GREEN)lint$(RESET)                - Run linters on the project."
-	@echo "  $(GREEN)setup-config$(RESET)        - Setup the configuration for OpenHands by providing LLM API key,"
+	@echo "  $(GREEN)setup-config$(RESET)        - Setup the configuration for OpenDevin by providing LLM API key,"
 	@echo "                        LLM Model name, and workspace directory."
-	@echo "  $(GREEN)start-backend$(RESET)       - Start the backend server for the OpenHands project."
-	@echo "  $(GREEN)start-frontend$(RESET)      - Start the frontend server for the OpenHands project."
-	@echo "  $(GREEN)run$(RESET)                 - Run the OpenHands application, starting both backend and frontend servers."
+	@echo "  $(GREEN)start-backend$(RESET)       - Start the backend server for the OpenDevin project."
+	@echo "  $(GREEN)start-frontend$(RESET)      - Start the frontend server for the OpenDevin project."
+	@echo "  $(GREEN)run$(RESET)                 - Run the OpenDevin application, starting both backend and frontend servers."
 	@echo "                        Backend Log file will be stored in the 'logs' directory."
-	@echo "  $(GREEN)docker-dev$(RESET)          - Build and run the OpenHands application in Docker."
-	@echo "  $(GREEN)docker-run$(RESET)          - Run the OpenHands application, starting both backend and frontend servers in Docker."
 	@echo "  $(GREEN)help$(RESET)                - Display this help message, providing information on available targets."

 # Phony targets
-.PHONY: build check-dependencies check-system check-python check-npm check-nodejs check-docker check-poetry install-python-dependencies install-frontend-dependencies install-pre-commit-hooks lint-backend lint-frontend lint test-frontend test build-frontend start-backend start-frontend _run_setup run run-wsl setup-config setup-config-prompts setup-config-basic openhands-cloud-run docker-dev docker-run clean help
-.PHONY: kind
+.PHONY: build check-dependencies check-python check-npm check-docker check-poetry pull-docker-image install-python-dependencies install-frontend-dependencies install-pre-commit-hooks lint start-backend start-frontend run run-wsl setup-config setup-config-prompts help
--- a/README.md
+++ b/README.md
@@ -1,153 +1,149 @@
 <a name="readme-top"></a>

-<div align="center">
-  <img src="https://raw.githubusercontent.com/OpenHands/docs/main/openhands/static/img/logo.png" alt="Logo" width="200">
-  <h1 align="center" style="border-bottom: none">OpenHands: AI-Driven Development</h1>
-</div>
+<!--
+*** Thanks for checking out the Best-README-Template. If you have a suggestion
+*** that would make this better, please fork the repo and create a pull request
+*** or simply open an issue with the tag "enhancement".
+*** Don't forget to give the project a star!
+*** Thanks again! Now go create something AMAZING! :D
+-->

+<!-- PROJECT SHIELDS -->
+<!--
+*** I'm using markdown "reference style" links for readability.
+*** Reference links are enclosed in brackets [ ] instead of parentheses ( ).
+*** See the bottom of this document for the declaration of the reference variables
+*** for contributors-url, forks-url, etc. This is an optional, concise syntax you may use.
+*** https://www.markdownguide.org/basic-syntax/#reference-style-links
+-->

 <div align="center">
-  <a href="https://github.com/OpenHands/OpenHands/blob/main/LICENSE"><img src="https://img.shields.io/badge/LICENSE-MIT-20B2AA?style=for-the-badge" alt="MIT License"></a>
-  <a href="https://docs.google.com/spreadsheets/d/1wOUdFCMyY6Nt0AIqF705KN4JKOWgeI4wUGUP60krXXs/edit?gid=811504672#gid=811504672"><img src="https://img.shields.io/badge/SWEBench-77.6-00cc00?logoColor=FFE165&style=for-the-badge" alt="Benchmark Score"></a>
+  <a href="https://github.com/OpenDevin/OpenDevin/graphs/contributors"><img src="https://img.shields.io/github/contributors/opendevin/opendevin?style=for-the-badge&color=blue" alt="Contributors"></a>
+  <a href="https://github.com/OpenDevin/OpenDevin/network/members"><img src="https://img.shields.io/github/forks/opendevin/opendevin?style=for-the-badge&color=blue" alt="Forks"></a>
+  <a href="https://github.com/OpenDevin/OpenDevin/stargazers"><img src="https://img.shields.io/github/stars/opendevin/opendevin?style=for-the-badge&color=blue" alt="Stargazers"></a>
+  <a href="https://github.com/OpenDevin/OpenDevin/issues"><img src="https://img.shields.io/github/issues/opendevin/opendevin?style=for-the-badge&color=blue" alt="Issues"></a>
+  <a href="https://github.com/OpenDevin/OpenDevin/blob/main/LICENSE"><img src="https://img.shields.io/github/license/opendevin/opendevin?style=for-the-badge&color=blue" alt="MIT License"></a>
  <br/>
-  <a href="https://docs.openhands.dev/sdk"><img src="https://img.shields.io/badge/Documentation-000?logo=googledocs&logoColor=FFE165&style=for-the-badge" alt="Check out the documentation"></a>
-  <a href="https://arxiv.org/abs/2511.03690"><img src="https://img.shields.io/badge/Paper-000?logoColor=FFE165&logo=arxiv&style=for-the-badge" alt="Tech Report"></a>
-
-
-  <!-- Keep these links. Translations will automatically update with the README. -->
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=de">Deutsch</a> |
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=es">Español</a> |
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=fr">français</a> |
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=ja">日本語</a> |
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=ko">한국어</a> |
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=pt">Português</a> |
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=ru">Русский</a> |
-  <a href="https://www.readme-i18n.com/OpenHands/OpenHands?lang=zh">中文</a>
+  <a href="https://join.slack.com/t/opendevin/shared_invite/zt-2i1iqdag6-bVmvamiPA9EZUu7oCO6KhA"><img src="https://img.shields.io/badge/Slack-Join%20Us-red?logo=slack&logoColor=white&style=for-the-badge" alt="Join our Slack community"></a>
+  <a href="https://discord.gg/ESHStjSjD4"><img src="https://img.shields.io/badge/Discord-Join%20Us-purple?logo=discord&logoColor=white&style=for-the-badge" alt="Join our Discord community"></a>
+  <a href="https://codecov.io/github/opendevin/opendevin?branch=main"><img alt="CodeCov" src="https://img.shields.io/codecov/c/github/opendevin/opendevin?style=for-the-badge"></a>
 </div>

-<hr>
-
-🙌 Welcome to OpenHands, a [community](COMMUNITY.md) focused on AI-driven development. We’d love for you to [join us on Slack](https://dub.sh/openhands).
-
-There are a few ways to work with OpenHands:
-
-### OpenHands Software Agent SDK
-The SDK is a composable Python library that contains all of our agentic tech. It's the engine that powers everything else below.
-
-Define agents in code, then run them locally, or scale to 1000s of agents in the cloud.
-
-[Check out the docs](https://docs.openhands.dev/sdk) or [view the source](https://github.com/OpenHands/software-agent-sdk/)
-
-### OpenHands CLI
-The CLI is the easiest way to start using OpenHands. The experience will be familiar to anyone who has worked
-with e.g. Claude Code or Codex. You can power it with Claude, GPT, or any other LLM.
-
-[Check out the docs](https://docs.openhands.dev/openhands/usage/run-openhands/cli-mode) or [view the source](https://github.com/OpenHands/OpenHands-CLI)
-
-### OpenHands Local GUI
-Use the Local GUI for running agents on your laptop. It comes with a REST API and a single-page React application.
-The experience will be familiar to anyone who has used Devin or Jules.
-
-[Check out the docs](https://docs.openhands.dev/openhands/usage/run-openhands/local-setup) or view the source in this repo.
-
-### OpenHands Cloud
-This is a deployment of OpenHands GUI, running on hosted infrastructure.
-
-You can try it for free using the Minimax model by [signing in with your GitHub or GitLab account](https://app.all-hands.dev).
-
-OpenHands Cloud comes with source-available features and integrations:
- Integrations with Slack, Jira, and Linear
- Multi-user support
- RBAC and permissions
- Collaboration features (e.g., conversation sharing)
-
-### OpenHands Enterprise
-Large enterprises can work with us to self-host OpenHands Cloud in their own VPC, via Kubernetes.
-OpenHands Enterprise can also work with the CLI and SDK above.
-
-OpenHands Enterprise is source-available--you can see all the source code here in the enterprise/ directory,
-but you'll need to purchase a license if you want to run it for more than one month.
-
-Enterprise contracts also come with extended support and access to our research team.
-
-Learn more at [openhands.dev/enterprise](https://openhands.dev/enterprise)
-
-### Everything Else
-
-Check out our [Product Roadmap](https://github.com/orgs/openhands/projects/1), and feel free to
-[open up an issue](https://github.com/OpenHands/OpenHands/issues) if there's something you'd like to see!
-
-You might also be interested in our [evaluation infrastructure](https://github.com/OpenHands/benchmarks), our [chrome extension](https://github.com/OpenHands/openhands-chrome-extension/), or our [Theory-of-Mind module](https://github.com/OpenHands/ToM-SWE).
-
-All our work is available under the MIT license, except for the `enterprise/` directory in this repository (see the [enterprise license](enterprise/LICENSE) for details).
-The core `openhands` and `agent-server` Docker images are fully MIT-licensed as well.
-
-If you need help with anything, or just want to chat, [come find us on Slack](https://dub.sh/openhands).
-
-<hr>
-
-### Thank You to Our Contributors
-
+<!-- PROJECT LOGO -->
 <div align="center">
-
-[![OpenHands Contributors](https://assets.openhands.dev/readme/openhands-openhands-contributors.svg)](https://github.com/OpenHands/OpenHands/graphs/contributors)
-
+  <img src="./docs/static/img/logo.png" alt="Logo" width="200" height="200">
+  <h1 align="center">OpenDevin: Code Less, Make More</h1>
+  <a href="https://opendevin.github.io/OpenDevin/modules/usage/intro"><img src="https://img.shields.io/badge/Documentation-OpenDevin-blue?logo=googledocs&logoColor=white&style=for-the-badge" alt="Check out the documentation"></a>
+  <a href="https://huggingface.co/spaces/OpenDevin/evaluation"><img src="https://img.shields.io/badge/Evaluation-Benchmark%20on%20HF%20Space-green?style=for-the-badge" alt="Evaluation Benchmark"></a>
 </div>
-
 <hr>

-### Trusted by Engineers at
+Welcome to OpenDevin, a platform for autonomous software engineers, powered by AI and LLMs.

-<div align="center">
-  <br/><br/>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/tiktok.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/tiktok.svg" alt="TikTok" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/vmware.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/vmware.svg" alt="VMware" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/roche.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/roche.svg" alt="Roche" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/amazon.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/amazon.svg" alt="Amazon" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/c3-ai.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/c3-ai.svg" alt="C3 AI" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/netflix.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/netflix.svg" alt="Netflix" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/mastercard.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/mastercard.svg" alt="Mastercard" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/red-hat.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/red-hat.svg" alt="Red Hat" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/mongodb.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/mongodb.svg" alt="MongoDB" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/apple.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/apple.svg" alt="Apple" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/nvidia.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/nvidia.svg" alt="NVIDIA" height="17" hspace="5">
-  </picture>
-  <picture>
-    <source media="(prefers-color-scheme: dark)" srcset="https://assets.openhands.dev/logos/external/white/google.svg">
-    <img src="https://assets.openhands.dev/logos/external/black/google.svg" alt="Google" height="17" hspace="5">
-  </picture>
-</div>
+OpenDevin agents collaborate with human developers to write code, fix bugs, and ship features.

-</div>
+![App screenshot](./docs/static/img/screenshot.png)
+
+## ⚡ Getting Started
+OpenDevin works best with the most recent version of Docker, `26.0.0`.
+You must be using Linux, Mac OS, or WSL on Windows.
+
+To start OpenDevin in a docker container, run the following commands in your terminal:
+
+> [!WARNING]
+> When you run the following command, files in `./workspace` may be modified or deleted.
+
+```bash
+WORKSPACE_BASE=$(pwd)/workspace
+docker run -it \
+    --pull=always \
+    -e SANDBOX_USER_ID=$(id -u) \
+    -e WORKSPACE_MOUNT_PATH=$WORKSPACE_BASE \
+    -v $WORKSPACE_BASE:/opt/workspace_base \
+    -v /var/run/docker.sock:/var/run/docker.sock \
+    -p 3000:3000 \
+    --add-host host.docker.internal:host-gateway \
+    --name opendevin-app-$(date +%Y%m%d%H%M%S) \
+    ghcr.io/opendevin/opendevin
+```
+
+> [!NOTE]
+> By default, this command pulls the `latest` tag, which represents the most recent release of OpenDevin. You have other options as well:
+> - For a specific release version, use `ghcr.io/opendevin/opendevin:<OpenDevin_version>` (replace <OpenDevin_version> with the desired version number).
+> - For the most up-to-date development version, use `ghcr.io/opendevin/opendevin:main`. This version may be **(unstable!)** and is recommended for testing or development purposes only.
+> 
+> Choose the tag that best suits your needs based on stability requirements and desired features.
+
+You'll find OpenDevin running at [http://localhost:3000](http://localhost:3000) with access to `./workspace`. To have OpenDevin operate on your code, place it in `./workspace`.
+OpenDevin will only have access to this workspace folder. The rest of your system will not be affected as it runs in a secured docker sandbox.
+
+Upon opening OpenDevin, you must select the appropriate `Model` and enter the `API Key` within the settings that should pop up automatically. These can be set at any time by selecting
+the `Settings` button (gear icon) in the UI. If the required `Model` does not exist in the list, you can manually enter it in the text box.
+
+For the development workflow, see [Development.md](https://github.com/OpenDevin/OpenDevin/blob/main/Development.md).
+
+Are you having trouble? Check out our [Troubleshooting Guide](https://opendevin.github.io/OpenDevin/modules/usage/troubleshooting).
+
+## 🚀 Documentation
+
+To learn more about the project, and for tips on using OpenDevin,
+**check out our [documentation](https://opendevin.github.io/OpenDevin/modules/usage/intro)**.
+
+There you'll find resources on how to use different LLM providers (like ollama and Anthropic's Claude),
+troubleshooting resources, and advanced configuration options.
+
+## 🤝 How to Contribute
+
+OpenDevin is a community-driven project, and we welcome contributions from everyone.
+Whether you're a developer, a researcher, or simply enthusiastic about advancing the field of
+software engineering with AI, there are many ways to get involved:
+
+- **Code Contributions:** Help us develop new agents, core functionality, the frontend and other interfaces, or sandboxing solutions.
+- **Research and Evaluation:** Contribute to our understanding of LLMs in software engineering, participate in evaluating the models, or suggest improvements.
+- **Feedback and Testing:** Use the OpenDevin toolset, report bugs, suggest features, or provide feedback on usability.
+
+For details, please check [CONTRIBUTING.md](./CONTRIBUTING.md).
+
+## 🤖 Join Our Community
+
+Whether you're a developer, a researcher, or simply enthusiastic about OpenDevin, we'd love to have you in our community.
+Let's make software engineering better together!
+
+- [Slack workspace](https://join.slack.com/t/opendevin/shared_invite/zt-2jsrl32uf-fTeeFjNyNYxqSZt5NPY3fA) - Here we talk about research, architecture, and future development.
+- [Discord server](https://discord.gg/ESHStjSjD4) - This is a community-run server for general discussion, questions, and feedback.
+
+## 📈 Progress
+
+<p align="center">
+  <a href="https://star-history.com/#OpenDevin/OpenDevin&Date">
+    <img src="https://api.star-history.com/svg?repos=OpenDevin/OpenDevin&type=Date" width="500" alt="Star History Chart">
+  </a>
+</p>
+
+## 📜 License
+
+Distributed under the MIT License. See [`LICENSE`](./LICENSE) for more information.
+
+[contributors-shield]: https://img.shields.io/github/contributors/opendevin/opendevin?style=for-the-badge
+[contributors-url]: https://github.com/OpenDevin/OpenDevin/graphs/contributors
+[forks-shield]: https://img.shields.io/github/forks/opendevin/opendevin?style=for-the-badge
+[forks-url]: https://github.com/OpenDevin/OpenDevin/network/members
+[stars-shield]: https://img.shields.io/github/stars/opendevin/opendevin?style=for-the-badge
+[stars-url]: https://github.com/OpenDevin/OpenDevin/stargazers
+[issues-shield]: https://img.shields.io/github/issues/opendevin/opendevin?style=for-the-badge
+[issues-url]: https://github.com/OpenDevin/OpenDevin/issues
+[license-shield]: https://img.shields.io/github/license/opendevin/opendevin?style=for-the-badge
+[license-url]: https://github.com/OpenDevin/OpenDevin/blob/main/LICENSE
+
+## 📚 Cite
+
+```
+@misc{opendevin2024,
+  author       = {{OpenDevin Team}},
+  title        = {{OpenDevin: An Open Platform for AI Software Developers as Generalist Agents}},
+  year         = {2024},
+  version      = {v1.0},
+  howpublished = {\url{https://github.com/OpenDevin/OpenDevin}},
+  note         = {Accessed: ENTER THE DATE YOU ACCESSED THE PROJECT}
+}
+```
--- a/agenthub/README.md
+++ b/agenthub/README.md
@@ -0,0 +1,72 @@
+# Agent Framework Research
+
+In this folder, there may exist multiple implementations of `Agent` that will be used by the framework.
+
+For example, `agenthub/codeact_agent`, etc.
+Contributors from different backgrounds and interests can choose to contribute to any (or all!) of these directions.
+
+## Constructing an Agent
+
+The abstraction for an agent can be found [here](../opendevin/controller/agent.py).
+
+Agents are run inside of a loop. At each iteration, `agent.step()` is called with a
+[State](../opendevin/controller/state/state.py) input, and the agent must output an [Action](../opendevin/events/action).
+
+Every agent also has a `self.llm` which it can use to interact with the LLM configured by the user.
+See the [LiteLLM docs for `self.llm.completion`](https://docs.litellm.ai/docs/completion).
+
+## State
+
+The `state` contains:
+
+- A history of actions taken by the agent, as well as any observations (e.g. file content, command output) from those actions
+- A list of actions/observations that have happened since the most recent step
+- A [`root_task`](https://github.com/OpenDevin/OpenDevin/blob/main/opendevin/controller/state/task.py), which contains a plan of action
+  - The agent can add and modify subtasks through the `AddTaskAction` and `ModifyTaskAction`
+
+## Actions
+
+Here is a list of available Actions, which can be returned by `agent.step()`:
+
+- [`CmdRunAction`](../opendevin/events/action/commands.py) - Runs a command inside a sandboxed terminal
+- [`IPythonRunCellAction`](../opendevin/events/action/commands.py) - Execute a block of Python code interactively (in Jupyter notebook) and receives `CmdOutputObservation`. Requires setting up `jupyter` [plugin](../opendevin/runtime/plugins) as a requirement.
+- [`FileReadAction`](../opendevin/events/action/files.py) - Reads the content of a file
+- [`FileWriteAction`](../opendevin/events/action/files.py) - Writes new content to a file
+- [`BrowseURLAction`](../opendevin/events/action/browse.py) - Gets the content of a URL
+- [`AddTaskAction`](../opendevin/events/action/tasks.py) - Adds a subtask to the plan
+- [`ModifyTaskAction`](../opendevin/events/action/tasks.py) - Changes the state of a subtask.
+- [`AgentFinishAction`](../opendevin/events/action/agent.py) - Stops the control loop, allowing the user/delegator agent to enter a new task
+- [`AgentRejectAction`](../opendevin/events/action/agent.py) - Stops the control loop, allowing the user/delegator agent to enter a new task
+- [`AgentFinishAction`](../opendevin/events/action/agent.py) - Stops the control loop, allowing the user to enter a new task
+- [`MessageAction`](../opendevin/events/action/message.py) - Represents a message from an agent or the user
+
+You can use `action.to_dict()` and `action_from_dict` to serialize and deserialize actions.
+
+## Observations
+
+There are also several types of Observations. These are typically available in the step following the corresponding Action.
+But they may also appear as a result of asynchronous events (e.g. a message from the user).
+
+Here is a list of available Observations:
+
+- [`CmdOutputObservation`](../opendevin/events/observation/commands.py)
+- [`BrowserOutputObservation`](../opendevin/events/observation/browse.py)
+- [`FileReadObservation`](../opendevin/events/observation/files.py)
+- [`FileWriteObservation`](../opendevin/events/observation/files.py)
+- [`ErrorObservation`](../opendevin/events/observation/error.py)
+- [`SuccessObservation`](../opendevin/events/observation/success.py)
+
+You can use `observation.to_dict()` and `observation_from_dict` to serialize and deserialize observations.
+
+## Interface
+
+Every agent must implement the following methods:
+
+### `step`
+
+```
+def step(self, state: "State") -> "Action"
+```
+
+`step` moves the agent forward one step towards its goal. This probably means
+sending a prompt to the LLM, then parsing the response into an `Action`.
--- a/agenthub/init.py
+++ b/agenthub/init.py
@@ -0,0 +1,46 @@
+from dotenv import load_dotenv
+
+from opendevin.controller.agent import Agent
+
+from .micro.agent import MicroAgent
+from .micro.registry import all_microagents
+
+load_dotenv()
+
+
+from . import (  # noqa: E402
+    browsing_agent,
+    codeact_agent,
+    codeact_swe_agent,
+    delegator_agent,
+    dummy_agent,
+    gptswarm_agent,
+    monologue_agent,
+    planner_agent,
+)
+
+__all__ = [
+    'monologue_agent',
+    'codeact_agent',
+    'gptswarm_agent',
+    'codeact_swe_agent',
+    'planner_agent',
+    'delegator_agent',
+    'dummy_agent',
+    'browsing_agent',
+]
+
+for agent in all_microagents.values():
+    name = agent['name']
+    prompt = agent['prompt']
+
+    anon_class = type(
+        name,
+        (MicroAgent,),
+        {
+            'prompt': prompt,
+            'agent_definition': agent,
+        },
+    )
+
+    Agent.register(name, anon_class)
--- a/agenthub/browsing_agent/README.md
+++ b/agenthub/browsing_agent/README.md
@@ -0,0 +1,16 @@
+# Browsing Agent Framework
+
+This folder implements the basic BrowserGym [demo agent](https://github.com/ServiceNow/BrowserGym/tree/main/demo_agent) that enables full-featured web browsing.
+
+
+## Test run
+
+Note that for browsing tasks, GPT-4 is usually a requirement to get reasonable results, due to the complexity of the web page structures.
+
+```
+poetry run python ./opendevin/core/main.py \
+           -i 10 \
+           -t "tell me the usa's president using google search" \
+           -c BrowsingAgent \
+           -m gpt-4o-2024-05-13
+```
--- a/agenthub/browsing_agent/init.py
+++ b/agenthub/browsing_agent/init.py
@@ -0,0 +1,5 @@
+from opendevin.controller.agent import Agent
+
+from .browsing_agent import BrowsingAgent
+
+Agent.register('BrowsingAgent', BrowsingAgent)
--- a/agenthub/browsing_agent/browsing_agent.py
+++ b/agenthub/browsing_agent/browsing_agent.py
@@ -0,0 +1,215 @@
+import os
+
+from browsergym.core.action.highlevel import HighLevelActionSet
+from browsergym.utils.obs import flatten_axtree_to_str
+
+from agenthub.browsing_agent.response_parser import BrowsingResponseParser
+from opendevin.controller.agent import Agent
+from opendevin.controller.state.state import State
+from opendevin.core.logger import opendevin_logger as logger
+from opendevin.events.action import (
+    Action,
+    AgentFinishAction,
+    BrowseInteractiveAction,
+    MessageAction,
+)
+from opendevin.events.event import EventSource
+from opendevin.events.observation import BrowserOutputObservation
+from opendevin.events.observation.observation import Observation
+from opendevin.llm.llm import LLM
+from opendevin.runtime.plugins import (
+    PluginRequirement,
+)
+from opendevin.runtime.tools import RuntimeTool
+
+USE_NAV = (
+    os.environ.get('USE_NAV', 'true') == 'true'
+)  # only disable NAV actions when running webarena and miniwob benchmarks
+USE_CONCISE_ANSWER = (
+    os.environ.get('USE_CONCISE_ANSWER', 'false') == 'true'
+)  # only return concise answer when running webarena and miniwob benchmarks
+
+if not USE_NAV and USE_CONCISE_ANSWER:
+    EVAL_MODE = True  # disabled NAV actions and only return concise answer, for webarena and miniwob benchmarks\
+else:
+    EVAL_MODE = False
+
+
+def get_error_prefix(last_browser_action: str) -> str:
+    return f'IMPORTANT! Last action is incorrect:\n{last_browser_action}\nThink again with the current observation of the page.\n'
+
+
+def get_system_message(goal: str, action_space: str) -> str:
+    return f"""\
+# Instructions
+Review the current state of the page and all other information to find the best
+possible next action to accomplish your goal. Your answer will be interpreted
+and executed by a program, make sure to follow the formatting instructions.
+
+# Goal:
+{goal}
+
+# Action Space
+{action_space}
+"""
+
+
+CONCISE_INSTRUCTION = """\
+
+Here is another example with chain of thought of a valid action when providing a concise answer to user:
+"
+In order to accomplish my goal I need to send the information asked back to the user. This page list the information of HP Inkjet Fax Machine, which is the product identified in the objective. Its price is $279.49. I will send a message back to user with the answer.
+```send_msg_to_user("$279.49")```
+"
+"""
+
+
+def get_prompt(error_prefix: str, cur_axtree_txt: str, prev_action_str: str) -> str:
+    prompt = f"""\
+{error_prefix}
+
+# Current Accessibility Tree:
+{cur_axtree_txt}
+
+# Previous Actions
+{prev_action_str}
+
+Here is an example with chain of thought of a valid action when clicking on a button:
+"
+In order to accomplish my goal I need to click on the button with bid 12
+```click("12")```
+"
+""".strip()
+    if USE_CONCISE_ANSWER:
+        prompt += CONCISE_INSTRUCTION
+    return prompt
+
+
+class BrowsingAgent(Agent):
+    VERSION = '1.0'
+    """
+    An agent that interacts with the browser.
+    """
+
+    sandbox_plugins: list[PluginRequirement] = []
+    runtime_tools: list[RuntimeTool] = [RuntimeTool.BROWSER]
+    response_parser = BrowsingResponseParser()
+
+    def __init__(
+        self,
+        llm: LLM,
+    ) -> None:
+        """
+        Initializes a new instance of the BrowsingAgent class.
+
+        Parameters:
+        - llm (LLM): The llm to be used by this agent
+        """
+        super().__init__(llm)
+        # define a configurable action space, with chat functionality, web navigation, and webpage grounding using accessibility tree and HTML.
+        # see https://github.com/ServiceNow/BrowserGym/blob/main/core/src/browsergym/core/action/highlevel.py for more details
+        action_subsets = ['chat', 'bid']
+        if USE_NAV:
+            action_subsets.append('nav')
+        self.action_space = HighLevelActionSet(
+            subsets=action_subsets,
+            strict=False,  # less strict on the parsing of the actions
+            multiaction=True,  # enable to agent to take multiple actions at once
+        )
+
+        self.reset()
+
+    def reset(self) -> None:
+        """
+        Resets the Browsing Agent.
+        """
+        super().reset()
+        self.cost_accumulator = 0
+        self.error_accumulator = 0
+
+    def step(self, state: State) -> Action:
+        """
+        Performs one step using the Browsing Agent.
+        This includes gathering information on previous steps and prompting the model to make a browsing command to execute.
+
+        Parameters:
+        - state (State): used to get updated info
+
+        Returns:
+        - BrowseInteractiveAction(browsergym_command) - BrowserGym commands to run
+        - MessageAction(content) - Message action to run (e.g. ask for clarification)
+        - AgentFinishAction() - end the interaction
+        """
+        messages = []
+        prev_actions = []
+        cur_axtree_txt = ''
+        error_prefix = ''
+        last_obs = None
+        last_action = None
+
+        if EVAL_MODE and len(state.history.get_events_as_list()) == 1:
+            # for webarena and miniwob++ eval, we need to retrieve the initial observation already in browser env
+            # initialize and retrieve the first observation by issuing an noop OP
+            # For non-benchmark browsing, the browser env starts with a blank page, and the agent is expected to first navigate to desired websites
+            return BrowseInteractiveAction(browser_actions='noop()')
+
+        for event in state.history.get_events():
+            if isinstance(event, BrowseInteractiveAction):
+                prev_actions.append(event.browser_actions)
+                last_action = event
+            elif isinstance(event, MessageAction) and event.source == EventSource.AGENT:
+                # agent has responded, task finished.
+                return AgentFinishAction(outputs={'content': event.content})
+            elif isinstance(event, Observation):
+                last_obs = event
+
+        if EVAL_MODE:
+            prev_actions = prev_actions[1:]  # remove the first noop action
+
+        prev_action_str = '\n'.join(prev_actions)
+        # if the final BrowserInteractiveAction exec BrowserGym's send_msg_to_user,
+        # we should also send a message back to the user in OpenDevin and call it a day
+        if (
+            isinstance(last_action, BrowseInteractiveAction)
+            and last_action.browsergym_send_msg_to_user
+        ):
+            return MessageAction(last_action.browsergym_send_msg_to_user)
+
+        if isinstance(last_obs, BrowserOutputObservation):
+            if last_obs.error:
+                # add error recovery prompt prefix
+                error_prefix = get_error_prefix(last_obs.last_browser_action)
+                self.error_accumulator += 1
+                if self.error_accumulator > 5:
+                    return MessageAction('Too many errors encountered. Task failed.')
+            try:
+                cur_axtree_txt = flatten_axtree_to_str(
+                    last_obs.axtree_object,
+                    extra_properties=last_obs.extra_element_properties,
+                    with_clickable=True,
+                    filter_visible_only=True,
+                )
+            except Exception as e:
+                logger.error(
+                    'Error when trying to process the accessibility tree: %s', e
+                )
+                return MessageAction('Error encountered when browsing.')
+
+        if (goal := state.get_current_user_intent()) is None:
+            goal = state.inputs['task']
+        system_msg = get_system_message(
+            goal,
+            self.action_space.describe(with_long_description=False, with_examples=True),
+        )
+
+        messages.append({'role': 'system', 'content': system_msg})
+
+        prompt = get_prompt(error_prefix, cur_axtree_txt, prev_action_str)
+        messages.append({'role': 'user', 'content': prompt})
+        logger.debug(prompt)
+        response = self.llm.completion(
+            messages=messages,
+            temperature=0.0,
+            stop=[')```', ')\n```'],
+        )
+        return self.response_parser.parse(response)
--- a/agenthub/browsing_agent/prompt.py
+++ b/agenthub/browsing_agent/prompt.py
@@ -0,0 +1,787 @@
+import abc
+import difflib
+import logging
+import platform
+from copy import deepcopy
+from dataclasses import asdict, dataclass
+from textwrap import dedent
+from typing import Literal, Union
+from warnings import warn
+
+from browsergym.core.action.base import AbstractActionSet
+from browsergym.core.action.highlevel import HighLevelActionSet
+from browsergym.core.action.python import PythonActionSet
+
+from opendevin.runtime.browser.browser_env import BrowserEnv
+
+from .utils import (
+    ParseError,
+    parse_html_tags_raise,
+)
+
+
+@dataclass
+class Flags:
+    use_html: bool = True
+    use_ax_tree: bool = False
+    drop_ax_tree_first: bool = True  # This flag is no longer active TODO delete
+    use_thinking: bool = False
+    use_error_logs: bool = False
+    use_past_error_logs: bool = False
+    use_history: bool = False
+    use_action_history: bool = False
+    use_memory: bool = False
+    use_diff: bool = False
+    html_type: str = 'pruned_html'
+    use_concrete_example: bool = True
+    use_abstract_example: bool = False
+    multi_actions: bool = False
+    action_space: Literal[
+        'python', 'bid', 'coord', 'bid+coord', 'bid+nav', 'coord+nav', 'bid+coord+nav'
+    ] = 'bid'
+    is_strict: bool = False
+    # This flag will be automatically disabled `if not chat_model_args.has_vision()`
+    use_screenshot: bool = True
+    enable_chat: bool = False
+    max_prompt_tokens: int = 100_000
+    extract_visible_tag: bool = False
+    extract_coords: Literal['False', 'center', 'box'] = 'False'
+    extract_visible_elements_only: bool = False
+    demo_mode: Literal['off', 'default', 'only_visible_elements'] = 'off'
+
+    def copy(self):
+        return deepcopy(self)
+
+    def asdict(self):
+        """Helper for JSON serializble requirement."""
+        return asdict(self)
+
+    @classmethod
+    def from_dict(self, flags_dict):
+        """Helper for JSON serializble requirement."""
+        if isinstance(flags_dict, Flags):
+            return flags_dict
+
+        if not isinstance(flags_dict, dict):
+            raise ValueError(
+                f'Unregcognized type for flags_dict of type {type(flags_dict)}.'
+            )
+        return Flags(**flags_dict)
+
+
+class PromptElement:
+    """Base class for all prompt elements. Prompt elements can be hidden.
+
+    Prompt elements are used to build the prompt. Use flags to control which
+    prompt elements are visible. We use class attributes as a convenient way
+    to implement static prompts, but feel free to override them with instance
+    attributes or @property decorator."""
+
+    _prompt = ''
+    _abstract_ex = ''
+    _concrete_ex = ''
+
+    def __init__(self, visible: bool = True) -> None:
+        """Prompt element that can be hidden.
+
+        Parameters
+        ----------
+        visible : bool, optional
+            Whether the prompt element should be visible, by default True. Can
+            be a callable that returns a bool. This is useful when a specific
+            flag changes during a shrink iteration.
+        """
+        self._visible = visible
+
+    @property
+    def prompt(self):
+        """Avoid overriding this method. Override _prompt instead."""
+        return self._hide(self._prompt)
+
+    @property
+    def abstract_ex(self):
+        """Useful when this prompt element is requesting an answer from the llm.
+        Provide an abstract example of the answer here. See Memory for an
+        example.
+
+        Avoid overriding this method. Override _abstract_ex instead
+        """
+        return self._hide(self._abstract_ex)
+
+    @property
+    def concrete_ex(self):
+        """Useful when this prompt element is requesting an answer from the llm.
+        Provide a concrete example of the answer here. See Memory for an
+        example.
+
+        Avoid overriding this method. Override _concrete_ex instead
+        """
+        return self._hide(self._concrete_ex)
+
+    @property
+    def is_visible(self):
+        """Handle the case where visible is a callable."""
+        visible = self._visible
+        if callable(visible):
+            visible = visible()
+        return visible
+
+    def _hide(self, value):
+        """Return value if visible is True, else return empty string."""
+        if self.is_visible:
+            return value
+        else:
+            return ''
+
+    def _parse_answer(self, text_answer) -> dict:
+        if self.is_visible:
+            return self._parse_answer(text_answer)
+        else:
+            return {}
+
+
+class Shrinkable(PromptElement, abc.ABC):
+    @abc.abstractmethod
+    def shrink(self) -> None:
+        """Implement shrinking of this prompt element.
+
+        You need to recursively call all shrinkable elements that are part of
+        this prompt. You can also implement a shrinking strategy for this prompt.
+        Shrinking is can be called multiple times to progressively shrink the
+        prompt until it fits max_tokens. Default max shrink iterations is 20.
+        """
+        pass
+
+
+class Truncater(Shrinkable):
+    """A prompt element that can be truncated to fit the context length of the LLM.
+    Of course, it will be great that we never have to use the functionality here to `shrink()` the prompt.
+    Extend this class for prompt elements that can be truncated. Usually long observations such as AxTree or HTML.
+    """
+
+    def __init__(self, visible, shrink_speed=0.3, start_truncate_iteration=10):
+        super().__init__(visible=visible)
+        self.shrink_speed = shrink_speed  # the percentage shrunk in each iteration
+        self.start_truncate_iteration = (
+            start_truncate_iteration  # the iteration to start truncating
+        )
+        self.shrink_calls = 0
+        self.deleted_lines = 0
+
+    def shrink(self) -> None:
+        if self.is_visible and self.shrink_calls >= self.start_truncate_iteration:
+            # remove the fraction of _prompt
+            lines = self._prompt.splitlines()
+            new_line_count = int(len(lines) * (1 - self.shrink_speed))
+            self.deleted_lines += len(lines) - new_line_count
+            self._prompt = '\n'.join(lines[:new_line_count])
+            self._prompt += (
+                f'\n... Deleted {self.deleted_lines} lines to reduce prompt size.'
+            )
+
+        self.shrink_calls += 1
+
+
+def fit_tokens(
+    shrinkable: Shrinkable,
+    max_prompt_chars=None,
+    max_iterations=20,
+):
+    """Shrink a prompt element until it fits max_tokens.
+
+    Parameters
+    ----------
+    shrinkable : Shrinkable
+        The prompt element to shrink.
+    max_prompt_chars : int
+        The maximum number of chars allowed.
+    max_iterations : int, optional
+        The maximum number of shrink iterations, by default 20.
+    model_name : str, optional
+        The name of the model used when tokenizing.
+
+    Returns
+    -------
+    str : the prompt after shrinking.
+    """
+
+    if max_prompt_chars is None:
+        return shrinkable.prompt
+
+    for _ in range(max_iterations):
+        prompt = shrinkable.prompt
+        if isinstance(prompt, str):
+            prompt_str = prompt
+        elif isinstance(prompt, list):
+            prompt_str = '\n'.join([p['text'] for p in prompt if p['type'] == 'text'])
+        else:
+            raise ValueError(f'Unrecognized type for prompt: {type(prompt)}')
+        n_chars = len(prompt_str)
+        if n_chars <= max_prompt_chars:
+            return prompt
+        shrinkable.shrink()
+
+    logging.info(
+        dedent(
+            f"""\
+            After {max_iterations} shrink iterations, the prompt is still
+            {len(prompt_str)} chars (greater than {max_prompt_chars}). Returning the prompt as is."""
+        )
+    )
+    return prompt
+
+
+class HTML(Truncater):
+    def __init__(self, html, visible: bool = True, prefix='') -> None:
+        super().__init__(visible=visible, start_truncate_iteration=5)
+        self._prompt = f'\n{prefix}HTML:\n{html}\n'
+
+
+class AXTree(Truncater):
+    def __init__(
+        self, ax_tree, visible: bool = True, coord_type=None, prefix=''
+    ) -> None:
+        super().__init__(visible=visible, start_truncate_iteration=10)
+        if coord_type == 'center':
+            coord_note = """\
+Note: center coordinates are provided in parenthesis and are
+  relative to the top left corner of the page.\n\n"""
+        elif coord_type == 'box':
+            coord_note = """\
+Note: bounding box of each object are provided in parenthesis and are
+  relative to the top left corner of the page.\n\n"""
+        else:
+            coord_note = ''
+        self._prompt = f'\n{prefix}AXTree:\n{coord_note}{ax_tree}\n'
+
+
+class Error(PromptElement):
+    def __init__(self, error, visible: bool = True, prefix='') -> None:
+        super().__init__(visible=visible)
+        self._prompt = f'\n{prefix}Error from previous action:\n{error}\n'
+
+
+class Observation(Shrinkable):
+    """Observation of the current step.
+
+    Contains the html, the accessibility tree and the error logs.
+    """
+
+    def __init__(self, obs, flags: Flags) -> None:
+        super().__init__()
+        self.flags = flags
+        self.obs = obs
+        self.html = HTML(obs[flags.html_type], visible=flags.use_html, prefix='## ')
+        self.ax_tree = AXTree(
+            obs['axtree_txt'],
+            visible=flags.use_ax_tree,
+            coord_type=flags.extract_coords,
+            prefix='## ',
+        )
+        self.error = Error(
+            obs['last_action_error'],
+            visible=flags.use_error_logs and obs['last_action_error'],
+            prefix='## ',
+        )
+
+    def shrink(self):
+        self.ax_tree.shrink()
+        self.html.shrink()
+
+    @property
+    def _prompt(self) -> str:  # type: ignore
+        return f'\n# Observation of current step:\n{self.html.prompt}{self.ax_tree.prompt}{self.error.prompt}\n\n'
+
+    def add_screenshot(self, prompt):
+        if self.flags.use_screenshot:
+            if isinstance(prompt, str):
+                prompt = [{'type': 'text', 'text': prompt}]
+            img_url = BrowserEnv.image_to_jpg_base64_url(
+                self.obs['screenshot'], add_data_prefix=True
+            )
+            prompt.append({'type': 'image_url', 'image_url': img_url})
+
+        return prompt
+
+
+class MacNote(PromptElement):
+    def __init__(self) -> None:
+        super().__init__(visible=platform.system() == 'Darwin')
+        self._prompt = '\nNote: you are on mac so you should use Meta instead of Control for Control+C etc.\n'
+
+
+class BeCautious(PromptElement):
+    def __init__(self, visible: bool = True) -> None:
+        super().__init__(visible=visible)
+        self._prompt = """\
+\nBe very cautious. Avoid submitting anything before verifying the effect of your
+actions. Take the time to explore the effect of safe actions first. For example
+you can fill a few elements of a form, but don't click submit before verifying
+that everything was filled correctly.\n"""
+
+
+class GoalInstructions(PromptElement):
+    def __init__(self, goal, visible: bool = True) -> None:
+        super().__init__(visible)
+        self._prompt = f"""\
+# Instructions
+Review the current state of the page and all other information to find the best
+possible next action to accomplish your goal. Your answer will be interpreted
+and executed by a program, make sure to follow the formatting instructions.
+
+## Goal:
+{goal}
+"""
+
+
+class ChatInstructions(PromptElement):
+    def __init__(self, chat_messages, visible: bool = True) -> None:
+        super().__init__(visible)
+        self._prompt = """\
+# Instructions
+
+You are a UI Assistant, your goal is to help the user perform tasks using a web browser. You can
+communicate with the user via a chat, in which the user gives you instructions and in which you
+can send back messages. You have access to a web browser that both you and the user can see,
+and with which only you can interact via specific commands.
+
+Review the instructions from the user, the current state of the page and all other information
+to find the best possible next action to accomplish your goal. Your answer will be interpreted
+and executed by a program, make sure to follow the formatting instructions.
+
+## Chat messages:
+
+"""
+        self._prompt += '\n'.join(
+            [
+                f"""\
+ - [{msg['role']}] {msg['message']}"""
+                for msg in chat_messages
+            ]
+        )
+
+
+class SystemPrompt(PromptElement):
+    _prompt = """\
+You are an agent trying to solve a web task based on the content of the page and
+a user instructions. You can interact with the page and explore. Each time you
+submit an action it will be sent to the browser and you will receive a new page."""
+
+
+class MainPrompt(Shrinkable):
+    def __init__(
+        self,
+        obs_history,
+        actions,
+        memories,
+        thoughts,
+        flags: Flags,
+    ) -> None:
+        super().__init__()
+        self.flags = flags
+        self.history = History(obs_history, actions, memories, thoughts, flags)
+        if self.flags.enable_chat:
+            self.instructions: Union[ChatInstructions, GoalInstructions] = (
+                ChatInstructions(obs_history[-1]['chat_messages'])
+            )
+        else:
+            if (
+                'chat_messages' in obs_history[-1]
+                and sum(
+                    [msg['role'] == 'user' for msg in obs_history[-1]['chat_messages']]
+                )
+                > 1
+            ):
+                logging.warning(
+                    'Agent is in goal mode, but multiple user messages are present in the chat. Consider switching to `enable_chat=True`.'
+                )
+            self.instructions = GoalInstructions(obs_history[-1]['goal'])
+
+        self.obs = Observation(obs_history[-1], self.flags)
+        self.action_space = ActionSpace(self.flags)
+
+        self.think = Think(visible=flags.use_thinking)
+        self.memory = Memory(visible=flags.use_memory)
+
+    @property
+    def _prompt(self) -> str:  # type: ignore
+        prompt = f"""\
+{self.instructions.prompt}\
+{self.obs.prompt}\
+{self.history.prompt}\
+{self.action_space.prompt}\
+{self.think.prompt}\
+{self.memory.prompt}\
+"""
+
+        if self.flags.use_abstract_example:
+            prompt += f"""
+# Abstract Example
+
+Here is an abstract version of the answer with description of the content of
+each tag. Make sure you follow this structure, but replace the content with your
+answer:
+{self.think.abstract_ex}\
+{self.memory.abstract_ex}\
+{self.action_space.abstract_ex}\
+"""
+
+        if self.flags.use_concrete_example:
+            prompt += f"""
+# Concrete Example
+
+Here is a concrete example of how to format your answer.
+Make sure to follow the template with proper tags:
+{self.think.concrete_ex}\
+{self.memory.concrete_ex}\
+{self.action_space.concrete_ex}\
+"""
+        return self.obs.add_screenshot(prompt)
+
+    def shrink(self):
+        self.history.shrink()
+        self.obs.shrink()
+
+    def _parse_answer(self, text_answer):
+        ans_dict = {}
+        ans_dict.update(self.think._parse_answer(text_answer))
+        ans_dict.update(self.memory._parse_answer(text_answer))
+        ans_dict.update(self.action_space._parse_answer(text_answer))
+        return ans_dict
+
+
+class ActionSpace(PromptElement):
+    def __init__(self, flags: Flags) -> None:
+        super().__init__()
+        self.flags = flags
+        self.action_space = _get_action_space(flags)
+
+        self._prompt = (
+            f'# Action space:\n{self.action_space.describe()}{MacNote().prompt}\n'
+        )
+        self._abstract_ex = f"""
+<action>
+{self.action_space.example_action(abstract=True)}
+</action>
+"""
+        self._concrete_ex = f"""
+<action>
+{self.action_space.example_action(abstract=False)}
+</action>
+"""
+
+    def _parse_answer(self, text_answer):
+        ans_dict = parse_html_tags_raise(
+            text_answer, keys=['action'], merge_multiple=True
+        )
+
+        try:
+            # just check if action can be mapped to python code but keep action as is
+            # the environment will be responsible for mapping it to python
+            self.action_space.to_python_code(ans_dict['action'])
+        except Exception as e:
+            raise ParseError(
+                f'Error while parsing action\n: {e}\n'
+                'Make sure your answer is restricted to the allowed actions.'
+            )
+
+        return ans_dict
+
+
+def _get_action_space(flags: Flags) -> AbstractActionSet:
+    match flags.action_space:
+        case 'python':
+            action_space = PythonActionSet(strict=flags.is_strict)
+            if flags.multi_actions:
+                warn(
+                    f'Flag action_space={repr(flags.action_space)} incompatible with multi_actions={repr(flags.multi_actions)}.',
+                    stacklevel=2,
+                )
+            if flags.demo_mode != 'off':
+                warn(
+                    f'Flag action_space={repr(flags.action_space)} incompatible with demo_mode={repr(flags.demo_mode)}.',
+                    stacklevel=2,
+                )
+            return action_space
+        case 'bid':
+            action_subsets = ['chat', 'bid']
+        case 'coord':
+            action_subsets = ['chat', 'coord']
+        case 'bid+coord':
+            action_subsets = ['chat', 'bid', 'coord']
+        case 'bid+nav':
+            action_subsets = ['chat', 'bid', 'nav']
+        case 'coord+nav':
+            action_subsets = ['chat', 'coord', 'nav']
+        case 'bid+coord+nav':
+            action_subsets = ['chat', 'bid', 'coord', 'nav']
+        case _:
+            raise NotImplementedError(
+                f'Unknown action_space {repr(flags.action_space)}'
+            )
+
+    action_space = HighLevelActionSet(
+        subsets=action_subsets,
+        multiaction=flags.multi_actions,
+        strict=flags.is_strict,
+        demo_mode=flags.demo_mode,
+    )
+
+    return action_space
+
+
+class Memory(PromptElement):
+    _prompt = ''  # provided in the abstract and concrete examples
+
+    _abstract_ex = """
+<memory>
+Write down anything you need to remember for next steps. You will be presented
+with the list of previous memories and past actions.
+</memory>
+"""
+
+    _concrete_ex = """
+<memory>
+I clicked on bid 32 to activate tab 2. The accessibility tree should mention
+focusable for elements of the form at next step.
+</memory>
+"""
+
+    def _parse_answer(self, text_answer):
+        return parse_html_tags_raise(
+            text_answer, optional_keys=['memory'], merge_multiple=True
+        )
+
+
+class Think(PromptElement):
+    _prompt = ''
+
+    _abstract_ex = """
+<think>
+Think step by step. If you need to make calculations such as coordinates, write them here. Describe the effect
+that your previous action had on the current content of the page.
+</think>
+"""
+    _concrete_ex = """
+<think>
+My memory says that I filled the first name and last name, but I can't see any
+content in the form. I need to explore different ways to fill the form. Perhaps
+the form is not visible yet or some fields are disabled. I need to replan.
+</think>
+"""
+
+    def _parse_answer(self, text_answer):
+        return parse_html_tags_raise(
+            text_answer, optional_keys=['think'], merge_multiple=True
+        )
+
+
+def diff(previous, new):
+    """Return a string showing the difference between original and new.
+
+    If the difference is above diff_threshold, return the diff string."""
+
+    if previous == new:
+        return 'Identical', []
+
+    if len(previous) == 0 or previous is None:
+        return 'previous is empty', []
+
+    diff_gen = difflib.ndiff(previous.splitlines(), new.splitlines())
+
+    diff_lines = []
+    plus_count = 0
+    minus_count = 0
+    for line in diff_gen:
+        if line.strip().startswith('+'):
+            diff_lines.append(line)
+            plus_count += 1
+        elif line.strip().startswith('-'):
+            diff_lines.append(line)
+            minus_count += 1
+        else:
+            continue
+
+    header = f'{plus_count} lines added and {minus_count} lines removed:'
+
+    return header, diff_lines
+
+
+class Diff(Shrinkable):
+    def __init__(
+        self, previous, new, prefix='', max_line_diff=20, shrink_speed=2, visible=True
+    ) -> None:
+        super().__init__(visible=visible)
+        self.max_line_diff = max_line_diff
+        self.header, self.diff_lines = diff(previous, new)
+        self.shrink_speed = shrink_speed
+        self.prefix = prefix
+
+    def shrink(self):
+        self.max_line_diff -= self.shrink_speed
+        self.max_line_diff = max(1, self.max_line_diff)
+
+    @property
+    def _prompt(self) -> str:  # type: ignore
+        diff_str = '\n'.join(self.diff_lines[: self.max_line_diff])
+        if len(self.diff_lines) > self.max_line_diff:
+            original_count = len(self.diff_lines)
+            diff_str = f'{diff_str}\nDiff truncated, {original_count - self.max_line_diff} changes now shown.'
+        return f'{self.prefix}{self.header}\n{diff_str}\n'
+
+
+class HistoryStep(Shrinkable):
+    def __init__(
+        self, previous_obs, current_obs, action, memory, flags: Flags, shrink_speed=1
+    ) -> None:
+        super().__init__()
+        self.html_diff = Diff(
+            previous_obs[flags.html_type],
+            current_obs[flags.html_type],
+            prefix='\n### HTML diff:\n',
+            shrink_speed=shrink_speed,
+            visible=lambda: flags.use_html and flags.use_diff,
+        )
+        self.ax_tree_diff = Diff(
+            previous_obs['axtree_txt'],
+            current_obs['axtree_txt'],
+            prefix='\n### Accessibility tree diff:\n',
+            shrink_speed=shrink_speed,
+            visible=lambda: flags.use_ax_tree and flags.use_diff,
+        )
+        self.error = Error(
+            current_obs['last_action_error'],
+            visible=(
+                flags.use_error_logs
+                and current_obs['last_action_error']
+                and flags.use_past_error_logs
+            ),
+            prefix='### ',
+        )
+        self.shrink_speed = shrink_speed
+        self.action = action
+        self.memory = memory
+        self.flags = flags
+
+    def shrink(self):
+        super().shrink()
+        self.html_diff.shrink()
+        self.ax_tree_diff.shrink()
+
+    @property
+    def _prompt(self) -> str:  # type: ignore
+        prompt = ''
+
+        if self.flags.use_action_history:
+            prompt += f'\n### Action:\n{self.action}\n'
+
+        prompt += (
+            f'{self.error.prompt}{self.html_diff.prompt}{self.ax_tree_diff.prompt}'
+        )
+
+        if self.flags.use_memory and self.memory is not None:
+            prompt += f'\n### Memory:\n{self.memory}\n'
+
+        return prompt
+
+
+class History(Shrinkable):
+    def __init__(
+        self, history_obs, actions, memories, thoughts, flags: Flags, shrink_speed=1
+    ) -> None:
+        super().__init__(visible=flags.use_history)
+        assert len(history_obs) == len(actions) + 1
+        assert len(history_obs) == len(memories) + 1
+
+        self.shrink_speed = shrink_speed
+        self.history_steps: list[HistoryStep] = []
+
+        for i in range(1, len(history_obs)):
+            self.history_steps.append(
+                HistoryStep(
+                    history_obs[i - 1],
+                    history_obs[i],
+                    actions[i - 1],
+                    memories[i - 1],
+                    flags,
+                )
+            )
+
+    def shrink(self):
+        """Shrink individual steps"""
+        # TODO set the shrink speed of older steps to be higher
+        super().shrink()
+        for step in self.history_steps:
+            step.shrink()
+
+    @property
+    def _prompt(self):
+        prompts = ['# History of interaction with the task:\n']
+        for i, step in enumerate(self.history_steps):
+            prompts.append(f'## step {i}')
+            prompts.append(step.prompt)
+        return '\n'.join(prompts) + '\n'
+
+
+if __name__ == '__main__':
+    html_template = """
+    <html>
+    <body>
+    <div>
+    Hello World.
+    Step {}.
+    </div>
+    </body>
+    </html>
+    """
+
+    OBS_HISTORY = [
+        {
+            'goal': 'do this and that',
+            'pruned_html': html_template.format(1),
+            'axtree_txt': '[1] Click me',
+            'last_action_error': '',
+        },
+        {
+            'goal': 'do this and that',
+            'pruned_html': html_template.format(2),
+            'axtree_txt': '[1] Click me',
+            'last_action_error': '',
+        },
+        {
+            'goal': 'do this and that',
+            'pruned_html': html_template.format(3),
+            'axtree_txt': '[1] Click me',
+            'last_action_error': 'Hey, there is an error now',
+        },
+    ]
+    ACTIONS = ["click('41')", "click('42')"]
+    MEMORIES = ['memory A', 'memory B']
+    THOUGHTS = ['thought A', 'thought B']
+
+    flags = Flags(
+        use_html=True,
+        use_ax_tree=True,
+        use_thinking=True,
+        use_error_logs=True,
+        use_past_error_logs=True,
+        use_history=True,
+        use_action_history=True,
+        use_memory=True,
+        use_diff=True,
+        html_type='pruned_html',
+        use_concrete_example=True,
+        use_abstract_example=True,
+        use_screenshot=False,
+        multi_actions=True,
+    )
+
+    print(
+        MainPrompt(
+            obs_history=OBS_HISTORY,
+            actions=ACTIONS,
+            memories=MEMORIES,
+            thoughts=THOUGHTS,
+            flags=flags,
+        ).prompt
+    )
--- a/agenthub/browsing_agent/response_parser.py
+++ b/agenthub/browsing_agent/response_parser.py
@@ -0,0 +1,90 @@
+import ast
+
+from opendevin.controller.action_parser import ActionParser, ResponseParser
+from opendevin.core.logger import opendevin_logger as logger
+from opendevin.events.action import (
+    Action,
+    BrowseInteractiveAction,
+)
+
+
+class BrowsingResponseParser(ResponseParser):
+    def __init__(self):
+        # Need to pay attention to the item order in self.action_parsers
+        super().__init__()
+        self.action_parsers = [BrowsingActionParserMessage()]
+        self.default_parser = BrowsingActionParserBrowseInteractive()
+
+    def parse(self, response: str) -> Action:
+        action_str = self.parse_response(response)
+        return self.parse_action(action_str)
+
+    def parse_response(self, response) -> str:
+        action_str = response['choices'][0]['message']['content']
+        if action_str is None:
+            return ''
+        action_str = action_str.strip()
+        if not action_str.endswith('```'):
+            action_str = action_str + ')```'
+        logger.info(action_str)
+        return action_str
+
+    def parse_action(self, action_str: str) -> Action:
+        for action_parser in self.action_parsers:
+            if action_parser.check_condition(action_str):
+                return action_parser.parse(action_str)
+        return self.default_parser.parse(action_str)
+
+
+class BrowsingActionParserMessage(ActionParser):
+    """
+    Parser action:
+        - BrowseInteractiveAction(browser_actions) - unexpected response format, message back to user
+    """
+
+    def __init__(
+        self,
+    ):
+        pass
+
+    def check_condition(self, action_str: str) -> bool:
+        return '```' not in action_str
+
+    def parse(self, action_str: str) -> Action:
+        msg = f'send_msg_to_user("""{action_str}""")'
+        return BrowseInteractiveAction(
+            browser_actions=msg,
+            thought=action_str,
+            browsergym_send_msg_to_user=action_str,
+        )
+
+
+class BrowsingActionParserBrowseInteractive(ActionParser):
+    """
+    Parser action:
+        - BrowseInteractiveAction(browser_actions) - handle send message to user function call in BrowserGym
+    """
+
+    def __init__(
+        self,
+    ):
+        pass
+
+    def check_condition(self, action_str: str) -> bool:
+        return True
+
+    def parse(self, action_str: str) -> Action:
+        thought = action_str.split('```')[0].strip()
+        action_str = action_str.split('```')[1].strip()
+        msg_content = ''
+        for sub_action in action_str.split('\n'):
+            if 'send_msg_to_user(' in sub_action:
+                tree = ast.parse(sub_action)
+                args = tree.body[0].value.args  # type: ignore
+                msg_content = args[0].value
+
+        return BrowseInteractiveAction(
+            browser_actions=action_str,
+            thought=thought,
+            browsergym_send_msg_to_user=msg_content,
+        )
--- a/agenthub/browsing_agent/utils.py
+++ b/agenthub/browsing_agent/utils.py
@@ -0,0 +1,160 @@
+import collections
+import re
+from warnings import warn
+
+import yaml
+
+
+def yaml_parser(message):
+    """Parse a yaml message for the retry function."""
+
+    # saves gpt-3.5 from some yaml parsing errors
+    message = re.sub(r':\s*\n(?=\S|\n)', ': ', message)
+
+    try:
+        value = yaml.safe_load(message)
+        valid = True
+        retry_message = ''
+    except yaml.YAMLError as e:
+        warn(str(e), stacklevel=2)
+        value = {}
+        valid = False
+        retry_message = "Your response is not a valid yaml. Please try again and be careful to the format. Don't add any apology or comment, just the answer."
+    return value, valid, retry_message
+
+
+def _compress_chunks(text, identifier, skip_list, split_regex='\n\n+'):
+    """Compress a string by replacing redundant chunks by identifiers. Chunks are defined by the split_regex."""
+    text_list = re.split(split_regex, text)
+    text_list = [chunk.strip() for chunk in text_list]
+    counter = collections.Counter(text_list)
+    def_dict = {}
+    id = 0
+
+    # Store items that occur more than once in a dictionary
+    for item, count in counter.items():
+        if count > 1 and item not in skip_list and len(item) > 10:
+            def_dict[f'{identifier}-{id}'] = item
+            id += 1
+
+    # Replace redundant items with their identifiers in the text
+    compressed_text = '\n'.join(text_list)
+    for key, value in def_dict.items():
+        compressed_text = compressed_text.replace(value, key)
+
+    return def_dict, compressed_text
+
+
+def compress_string(text):
+    """Compress a string by replacing redundant paragraphs and lines with identifiers."""
+
+    # Perform paragraph-level compression
+    def_dict, compressed_text = _compress_chunks(
+        text, identifier='§', skip_list=[], split_regex='\n\n+'
+    )
+
+    # Perform line-level compression, skipping any paragraph identifiers
+    line_dict, compressed_text = _compress_chunks(
+        compressed_text, '¶', list(def_dict.keys()), split_regex='\n+'
+    )
+    def_dict.update(line_dict)
+
+    # Create a definitions section
+    def_lines = ['<definitions>']
+    for key, value in def_dict.items():
+        def_lines.append(f'{key}:\n{value}')
+    def_lines.append('</definitions>')
+    definitions = '\n'.join(def_lines)
+
+    return definitions + '\n' + compressed_text
+
+
+def extract_html_tags(text, keys):
+    """Extract the content within HTML tags for a list of keys.
+
+    Parameters
+    ----------
+    text : str
+        The input string containing the HTML tags.
+    keys : list of str
+        The HTML tags to extract the content from.
+
+    Returns
+    -------
+    dict
+        A dictionary mapping each key to a list of subset in `text` that match the key.
+
+    Notes
+    -----
+    All text and keys will be converted to lowercase before matching.
+
+    """
+    content_dict = {}
+    # text = text.lower()
+    # keys = set([k.lower() for k in keys])
+    for key in keys:
+        pattern = f'<{key}>(.*?)</{key}>'
+        matches = re.findall(pattern, text, re.DOTALL)
+        if matches:
+            content_dict[key] = [match.strip() for match in matches]
+    return content_dict
+
+
+class ParseError(Exception):
+    pass
+
+
+def parse_html_tags_raise(text, keys=(), optional_keys=(), merge_multiple=False):
+    """A version of parse_html_tags that raises an exception if the parsing is not successful."""
+    content_dict, valid, retry_message = parse_html_tags(
+        text, keys, optional_keys, merge_multiple=merge_multiple
+    )
+    if not valid:
+        raise ParseError(retry_message)
+    return content_dict
+
+
+def parse_html_tags(text, keys=(), optional_keys=(), merge_multiple=False):
+    """Satisfy the parse api, extracts 1 match per key and validates that all keys are present
+
+    Parameters
+    ----------
+    text : str
+        The input string containing the HTML tags.
+    keys : list of str
+        The HTML tags to extract the content from.
+    optional_keys : list of str
+        The HTML tags to extract the content from, but are optional.
+
+    Returns
+    -------
+    dict
+        A dictionary mapping each key to subset of `text` that match the key.
+    bool
+        Whether the parsing was successful.
+    str
+        A message to be displayed to the agent if the parsing was not successful.
+    """
+    all_keys = tuple(keys) + tuple(optional_keys)
+    content_dict = extract_html_tags(text, all_keys)
+    retry_messages = []
+
+    for key in all_keys:
+        if key not in content_dict:
+            if key not in optional_keys:
+                retry_messages.append(f'Missing the key <{key}> in the answer.')
+        else:
+            val = content_dict[key]
+            content_dict[key] = val[0]
+            if len(val) > 1:
+                if not merge_multiple:
+                    retry_messages.append(
+                        f'Found multiple instances of the key {key}. You should have only one of them.'
+                    )
+                else:
+                    # merge the multiple instances
+                    content_dict[key] = '\n'.join(val)
+
+    valid = len(retry_messages) == 0
+    retry_message = '\n'.join(retry_messages)
+    return content_dict, valid, retry_message
--- a/agenthub/codeact_agent/README.md
+++ b/agenthub/codeact_agent/README.md
@@ -0,0 +1,29 @@
+# CodeAct Agent Framework
+
+This folder implements the CodeAct idea ([paper](https://arxiv.org/abs/2402.01030), [tweet](https://twitter.com/xingyaow_/status/1754556835703751087)) that consolidates LLM agents’ **act**ions into a unified **code** action space for both *simplicity* and *performance* (see paper for more details).
+
+The conceptual idea is illustrated below. At each turn, the agent can:
+
+1. **Converse**: Communicate with humans in natural language to ask for clarification, confirmation, etc.
+2. **CodeAct**: Choose to perform the task by executing code
+   - Execute any valid Linux `bash` command
+   - Execute any valid `Python` code with [an interactive Python interpreter](https://ipython.org/). This is simulated through `bash` command, see plugin system below for more details.
+
+![image](https://github.com/OpenDevin/OpenDevin/assets/38853559/92b622e3-72ad-4a61-8f41-8c040b6d5fb3)
+
+## Plugin System
+
+To make the CodeAct agent more powerful with only access to `bash` action space, CodeAct agent leverages OpenDevin's plugin system:
+- [Jupyter plugin](https://github.com/OpenDevin/OpenDevin/tree/main/opendevin/runtime/plugins/jupyter): for IPython execution via bash command
+- [SWE-agent tool plugin](https://github.com/OpenDevin/OpenDevin/tree/main/opendevin/runtime/plugins/swe_agent_commands): Powerful bash command line tools for software development tasks introduced by [swe-agent](https://github.com/princeton-nlp/swe-agent).
+
+## Demo
+
+https://github.com/OpenDevin/OpenDevin/assets/38853559/f592a192-e86c-4f48-ad31-d69282d5f6ac
+
+*Example of CodeActAgent with `gpt-4-turbo-2024-04-09` performing a data science task (linear regression)*
+
+## Work-in-progress & Next step
+
+[] Support web-browsing
+[] Complete the workflow for CodeAct agent to submit Github PRs
--- a/agenthub/codeact_agent/init.py
+++ b/agenthub/codeact_agent/init.py
@@ -0,0 +1,5 @@
+from opendevin.controller.agent import Agent
+
+from .codeact_agent import CodeActAgent
+
+Agent.register('CodeActAgent', CodeActAgent)
--- a/agenthub/codeact_agent/action_parser.py
+++ b/agenthub/codeact_agent/action_parser.py
@@ -0,0 +1,183 @@
+import re
+
+from opendevin.controller.action_parser import ActionParser, ResponseParser
+from opendevin.events.action import (
+    Action,
+    AgentDelegateAction,
+    AgentFinishAction,
+    CmdRunAction,
+    IPythonRunCellAction,
+    MessageAction,
+)
+
+
+class CodeActResponseParser(ResponseParser):
+    """
+    Parser action:
+        - CmdRunAction(command) - bash command to run
+        - IPythonRunCellAction(code) - IPython code to run
+        - AgentDelegateAction(agent, inputs) - delegate action for (sub)task
+        - MessageAction(content) - Message action to run (e.g. ask for clarification)
+        - AgentFinishAction() - end the interaction
+    """
+
+    def __init__(self):
+        # Need pay attention to the item order in self.action_parsers
+        super().__init__()
+        self.action_parsers = [
+            CodeActActionParserFinish(),
+            CodeActActionParserCmdRun(),
+            CodeActActionParserIPythonRunCell(),
+            CodeActActionParserAgentDelegate(),
+        ]
+        self.default_parser = CodeActActionParserMessage()
+
+    def parse(self, response) -> Action:
+        action_str = self.parse_response(response)
+        return self.parse_action(action_str)
+
+    def parse_response(self, response) -> str:
+        action = response.choices[0].message.content
+        if action is None:
+            return ''
+        for lang in ['bash', 'ipython', 'browse']:
+            if f'<execute_{lang}>' in action and f'</execute_{lang}>' not in action:
+                action += f'</execute_{lang}>'
+        return action
+
+    def parse_action(self, action_str: str) -> Action:
+        for action_parser in self.action_parsers:
+            if action_parser.check_condition(action_str):
+                return action_parser.parse(action_str)
+        return self.default_parser.parse(action_str)
+
+
+class CodeActActionParserFinish(ActionParser):
+    """
+    Parser action:
+        - AgentFinishAction() - end the interaction
+    """
+
+    def __init__(
+        self,
+    ):
+        self.finish_command = None
+
+    def check_condition(self, action_str: str) -> bool:
+        self.finish_command = re.search(r'<finish>.*</finish>', action_str, re.DOTALL)
+        return self.finish_command is not None
+
+    def parse(self, action_str: str) -> Action:
+        assert (
+            self.finish_command is not None
+        ), 'self.finish_command should not be None when parse is called'
+        thought = action_str.replace(self.finish_command.group(0), '').strip()
+        return AgentFinishAction(thought=thought)
+
+
+class CodeActActionParserCmdRun(ActionParser):
+    """
+    Parser action:
+        - CmdRunAction(command) - bash command to run
+        - AgentFinishAction() - end the interaction
+    """
+
+    def __init__(
+        self,
+    ):
+        self.bash_command = None
+
+    def check_condition(self, action_str: str) -> bool:
+        self.bash_command = re.search(
+            r'<execute_bash>(.*?)</execute_bash>', action_str, re.DOTALL
+        )
+        return self.bash_command is not None
+
+    def parse(self, action_str: str) -> Action:
+        assert (
+            self.bash_command is not None
+        ), 'self.bash_command should not be None when parse is called'
+        thought = action_str.replace(self.bash_command.group(0), '').strip()
+        # a command was found
+        command_group = self.bash_command.group(1).strip()
+        if command_group.strip() == 'exit':
+            return AgentFinishAction()
+        return CmdRunAction(command=command_group, thought=thought)
+
+
+class CodeActActionParserIPythonRunCell(ActionParser):
+    """
+    Parser action:
+        - IPythonRunCellAction(code) - IPython code to run
+    """
+
+    def __init__(
+        self,
+    ):
+        self.python_code = None
+        self.jupyter_kernel_init_code: str = 'from agentskills import *'
+
+    def check_condition(self, action_str: str) -> bool:
+        self.python_code = re.search(
+            r'<execute_ipython>(.*?)</execute_ipython>', action_str, re.DOTALL
+        )
+        return self.python_code is not None
+
+    def parse(self, action_str: str) -> Action:
+        assert (
+            self.python_code is not None
+        ), 'self.python_code should not be None when parse is called'
+        code_group = self.python_code.group(1).strip()
+        thought = action_str.replace(self.python_code.group(0), '').strip()
+        return IPythonRunCellAction(
+            code=code_group,
+            thought=thought,
+            kernel_init_code=self.jupyter_kernel_init_code,
+        )
+
+
+class CodeActActionParserAgentDelegate(ActionParser):
+    """
+    Parser action:
+        - AgentDelegateAction(agent, inputs) - delegate action for (sub)task
+    """
+
+    def __init__(
+        self,
+    ):
+        self.agent_delegate = None
+
+    def check_condition(self, action_str: str) -> bool:
+        self.agent_delegate = re.search(
+            r'<execute_browse>(.*)</execute_browse>', action_str, re.DOTALL
+        )
+        return self.agent_delegate is not None
+
+    def parse(self, action_str: str) -> Action:
+        assert (
+            self.agent_delegate is not None
+        ), 'self.agent_delegate should not be None when parse is called'
+        thought = action_str.replace(self.agent_delegate.group(0), '').strip()
+        browse_actions = self.agent_delegate.group(1).strip()
+        task = f'{thought}. I should start with: {browse_actions}'
+        return AgentDelegateAction(agent='BrowsingAgent', inputs={'task': task})
+
+
+class CodeActActionParserMessage(ActionParser):
+    """
+    Parser action:
+        - MessageAction(content) - Message action to run (e.g. ask for clarification)
+    """
+
+    def __init__(
+        self,
+    ):
+        pass
+
+    def check_condition(self, action_str: str) -> bool:
+        # We assume the LLM is GOOD enough that when it returns pure natural language
+        # it wants to talk to the user
+        return True
+
+    def parse(self, action_str: str) -> Action:
+        return MessageAction(content=action_str, wait_for_response=True)
--- a/agenthub/codeact_agent/codeact_agent.py
+++ b/agenthub/codeact_agent/codeact_agent.py
@@ -0,0 +1,241 @@
+from agenthub.codeact_agent.action_parser import CodeActResponseParser
+from agenthub.codeact_agent.prompt import (
+    COMMAND_DOCS,
+    EXAMPLES,
+    GITHUB_MESSAGE,
+    SYSTEM_PREFIX,
+    SYSTEM_SUFFIX,
+)
+from opendevin.controller.agent import Agent
+from opendevin.controller.state.state import State
+from opendevin.core.config import config
+from opendevin.events.action import (
+    Action,
+    AgentDelegateAction,
+    AgentFinishAction,
+    CmdRunAction,
+    IPythonRunCellAction,
+    MessageAction,
+)
+from opendevin.events.observation import (
+    AgentDelegateObservation,
+    CmdOutputObservation,
+    IPythonRunCellObservation,
+)
+from opendevin.events.serialization.event import truncate_content
+from opendevin.llm.llm import LLM
+from opendevin.runtime.plugins import (
+    AgentSkillsRequirement,
+    JupyterRequirement,
+    PluginRequirement,
+)
+from opendevin.runtime.tools import RuntimeTool
+
+ENABLE_GITHUB = True
+
+
+def action_to_str(action: Action) -> str:
+    if isinstance(action, CmdRunAction):
+        return f'{action.thought}\n<execute_bash>\n{action.command}\n</execute_bash>'
+    elif isinstance(action, IPythonRunCellAction):
+        return f'{action.thought}\n<execute_ipython>\n{action.code}\n</execute_ipython>'
+    elif isinstance(action, AgentDelegateAction):
+        return f'{action.thought}\n<execute_browse>\n{action.inputs["task"]}\n</execute_browse>'
+    elif isinstance(action, MessageAction):
+        return action.content
+    return ''
+
+
+def get_action_message(action: Action) -> dict[str, str] | None:
+    if (
+        isinstance(action, AgentDelegateAction)
+        or isinstance(action, CmdRunAction)
+        or isinstance(action, IPythonRunCellAction)
+        or isinstance(action, MessageAction)
+    ):
+        return {
+            'role': 'user' if action.source == 'user' else 'assistant',
+            'content': action_to_str(action),
+        }
+    return None
+
+
+def get_observation_message(obs) -> dict[str, str] | None:
+    max_message_chars = config.get_llm_config_from_agent(
+        'CodeActAgent'
+    ).max_message_chars
+    if isinstance(obs, CmdOutputObservation):
+        content = 'OBSERVATION:\n' + truncate_content(obs.content, max_message_chars)
+        content += (
+            f'\n[Command {obs.command_id} finished with exit code {obs.exit_code}]'
+        )
+        return {'role': 'user', 'content': content}
+    elif isinstance(obs, IPythonRunCellObservation):
+        content = 'OBSERVATION:\n' + obs.content
+        # replace base64 images with a placeholder
+        splitted = content.split('\n')
+        for i, line in enumerate(splitted):
+            if '![image](data:image/png;base64,' in line:
+                splitted[i] = (
+                    '![image](data:image/png;base64, ...) already displayed to user'
+                )
+        content = '\n'.join(splitted)
+        content = truncate_content(content, max_message_chars)
+        return {'role': 'user', 'content': content}
+    elif isinstance(obs, AgentDelegateObservation):
+        content = 'OBSERVATION:\n' + truncate_content(
+            str(obs.outputs), max_message_chars
+        )
+        return {'role': 'user', 'content': content}
+    return None
+
+
+# FIXME: We can tweak these two settings to create MicroAgents specialized toward different area
+def get_system_message() -> str:
+    if ENABLE_GITHUB:
+        return f'{SYSTEM_PREFIX}\n{GITHUB_MESSAGE}\n\n{COMMAND_DOCS}\n\n{SYSTEM_SUFFIX}'
+    else:
+        return f'{SYSTEM_PREFIX}\n\n{COMMAND_DOCS}\n\n{SYSTEM_SUFFIX}'
+
+
+def get_in_context_example() -> str:
+    return EXAMPLES
+
+
+class CodeActAgent(Agent):
+    VERSION = '1.8'
+    """
+    The Code Act Agent is a minimalist agent.
+    The agent works by passing the model a list of action-observation pairs and prompting the model to take the next step.
+
+    ### Overview
+
+    This agent implements the CodeAct idea ([paper](https://arxiv.org/abs/2402.13463), [tweet](https://twitter.com/xingyaow_/status/1754556835703751087)) that consolidates LLM agents’ **act**ions into a unified **code** action space for both *simplicity* and *performance* (see paper for more details).
+
+    The conceptual idea is illustrated below. At each turn, the agent can:
+
+    1. **Converse**: Communicate with humans in natural language to ask for clarification, confirmation, etc.
+    2. **CodeAct**: Choose to perform the task by executing code
+    - Execute any valid Linux `bash` command
+    - Execute any valid `Python` code with [an interactive Python interpreter](https://ipython.org/). This is simulated through `bash` command, see plugin system below for more details.
+
+    ![image](https://github.com/OpenDevin/OpenDevin/assets/38853559/92b622e3-72ad-4a61-8f41-8c040b6d5fb3)
+
+    ### Plugin System
+
+    To make the CodeAct agent more powerful with only access to `bash` action space, CodeAct agent leverages OpenDevin's plugin system:
+    - [Jupyter plugin](https://github.com/OpenDevin/OpenDevin/tree/main/opendevin/runtime/plugins/jupyter): for IPython execution via bash command
+    - [SWE-agent tool plugin](https://github.com/OpenDevin/OpenDevin/tree/main/opendevin/runtime/plugins/swe_agent_commands): Powerful bash command line tools for software development tasks introduced by [swe-agent](https://github.com/princeton-nlp/swe-agent).
+
+    ### Demo
+
+    https://github.com/OpenDevin/OpenDevin/assets/38853559/f592a192-e86c-4f48-ad31-d69282d5f6ac
+
+    *Example of CodeActAgent with `gpt-4-turbo-2024-04-09` performing a data science task (linear regression)*
+
+    ### Work-in-progress & Next step
+
+    [] Support web-browsing
+    [] Complete the workflow for CodeAct agent to submit Github PRs
+
+    """
+
+    sandbox_plugins: list[PluginRequirement] = [
+        # NOTE: AgentSkillsRequirement need to go before JupyterRequirement, since
+        # AgentSkillsRequirement provides a lot of Python functions,
+        # and it needs to be initialized before Jupyter for Jupyter to use those functions.
+        AgentSkillsRequirement(),
+        JupyterRequirement(),
+    ]
+    runtime_tools: list[RuntimeTool] = [RuntimeTool.BROWSER]
+
+    system_message: str = get_system_message()
+    in_context_example: str = f"Here is an example of how you can interact with the environment for task solving:\n{get_in_context_example()}\n\nNOW, LET'S START!"
+
+    action_parser = CodeActResponseParser()
+
+    def __init__(
+        self,
+        llm: LLM,
+    ) -> None:
+        """
+        Initializes a new instance of the CodeActAgent class.
+
+        Parameters:
+        - llm (LLM): The llm to be used by this agent
+        """
+        super().__init__(llm)
+        self.reset()
+
+    def reset(self) -> None:
+        """
+        Resets the CodeAct Agent.
+        """
+        super().reset()
+
+    def step(self, state: State) -> Action:
+        """
+        Performs one step using the CodeAct Agent.
+        This includes gathering info on previous steps and prompting the model to make a command to execute.
+
+        Parameters:
+        - state (State): used to get updated info
+
+        Returns:
+        - CmdRunAction(command) - bash command to run
+        - IPythonRunCellAction(code) - IPython code to run
+        - AgentDelegateAction(agent, inputs) - delegate action for (sub)task
+        - MessageAction(content) - Message action to run (e.g. ask for clarification)
+        - AgentFinishAction() - end the interaction
+        """
+
+        # if we're done, go back
+        latest_user_message = state.history.get_last_user_message()
+        if latest_user_message and latest_user_message.strip() == '/exit':
+            return AgentFinishAction()
+
+        # prepare what we want to send to the LLM
+        messages: list[dict[str, str]] = self._get_messages(state)
+
+        response = self.llm.completion(
+            messages=messages,
+            stop=[
+                '</execute_ipython>',
+                '</execute_bash>',
+                '</execute_browse>',
+            ],
+            temperature=0.0,
+        )
+        return self.action_parser.parse(response)
+
+    def _get_messages(self, state: State) -> list[dict[str, str]]:
+        messages = [
+            {'role': 'system', 'content': self.system_message},
+            {'role': 'user', 'content': self.in_context_example},
+        ]
+
+        for event in state.history.get_events():
+            # create a regular message from an event
+            message = (
+                get_action_message(event)
+                if isinstance(event, Action)
+                else get_observation_message(event)
+            )
+
+            # add regular message
+            if message:
+                messages.append(message)
+
+        # the latest user message is important:
+        # we want to remind the agent of the environment constraints
+        latest_user_message = next(
+            (m for m in reversed(messages) if m['role'] == 'user'), None
+        )
+
+        # add a reminder to the prompt
+        if latest_user_message:
+            latest_user_message['content'] += (
+                f'\n\nENVIRONMENT REMINDER: You have {state.max_iterations - state.iteration} turns left to complete the task. When finished reply with <finish></finish>'
+            )
+
+        return messages
--- a/agenthub/codeact_agent/prompt.py
+++ b/agenthub/codeact_agent/prompt.py
@@ -0,0 +1,275 @@
+from opendevin.runtime.plugins import AgentSkillsRequirement
+
+_AGENT_SKILLS_DOCS = AgentSkillsRequirement.documentation
+
+COMMAND_DOCS = (
+    '\nApart from the standard Python library, the assistant can also use the following functions (already imported) in <execute_ipython> environment:\n'
+    f'{_AGENT_SKILLS_DOCS}'
+    "Please note that THE `edit_file_by_replace`, `append_file` and `insert_content_at_line` FUNCTIONS REQUIRE PROPER INDENTATION. If the assistant would like to add the line '        print(x)', it must fully write that out, with all those spaces before the code! Indentation is important and code that is not indented correctly will fail and require fixing before it can be run."
+)
+
+# ======= SYSTEM MESSAGE =======
+MINIMAL_SYSTEM_PREFIX = """A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions.
+The assistant can use an interactive Python (Jupyter Notebook) environment, executing code with <execute_ipython>.
+<execute_ipython>
+print("Hello World!")
+</execute_ipython>
+The assistant can execute bash commands on behalf of the user by wrapping them with <execute_bash> and </execute_bash>.
+
+For example, you can list the files in the current directory by <execute_bash> ls </execute_bash>.
+Important, however: do not run interactive commands. You do not have access to stdin.
+Also, you need to handle commands that may run indefinitely and not return a result. For such cases, you should redirect the output to a file and run the command in the background to avoid blocking the execution.
+For example, to run a Python script that might run indefinitely without returning immediately, you can use the following format: <execute_bash> python3 app.py > server.log 2>&1 & </execute_bash>
+Also, if a command execution result saying like: Command: "npm start" timed out. Sending SIGINT to the process, you should also retry with running the command in the background.
+"""
+
+BROWSING_PREFIX = """The assistant can browse the Internet with <execute_browse> and </execute_browse>.
+For example, <execute_browse> Tell me the usa's president using google search </execute_browse>.
+Or <execute_browse> Tell me what is in http://example.com </execute_browse>.
+"""
+PIP_INSTALL_PREFIX = """The assistant can install Python packages using the %pip magic command in an IPython environment by using the following syntax: <execute_ipython> %pip install [package needed] </execute_ipython> and should always import packages and define variables before starting to use them."""
+
+SYSTEM_PREFIX = MINIMAL_SYSTEM_PREFIX + BROWSING_PREFIX + PIP_INSTALL_PREFIX
+
+GITHUB_MESSAGE = """To interact with GitHub, use the $GITHUB_TOKEN environment variable.
+For example, to push a branch `my_branch` to the GitHub repo `owner/repo`:
+<execute_bash> git push https://$GITHUB_TOKEN@github.com/owner/repo.git my_branch </execute_bash>
+If $GITHUB_TOKEN is not set, ask the user to set it."""
+
+SYSTEM_SUFFIX = """Responses should be concise.
+The assistant should attempt fewer things at a time instead of putting too many commands OR too much code in one "execute" block.
+Include ONLY ONE <execute_ipython>, <execute_bash>, or <execute_browse> per response, unless the assistant is finished with the task or needs more input or action from the user in order to proceed.
+If the assistant is finished with the task you MUST include <finish></finish> in your response.
+IMPORTANT: Execute code using <execute_ipython>, <execute_bash>, or <execute_browse> whenever possible.
+When handling files, try to use full paths and pwd to avoid errors.
+"""
+
+
+# ======= EXAMPLE MESSAGE =======
+EXAMPLES = """
+--- START OF EXAMPLE ---
+
+USER: Create a list of numbers from 1 to 10, and display them in a web page at port 5000.
+
+ASSISTANT:
+Sure! Let me create a Python file `app.py`:
+<execute_ipython>
+create_file('app.py')
+</execute_ipython>
+
+USER:
+OBSERVATION:
+[File: /workspace/app.py (1 lines total)]
+(this is the beginning of the file)
+1|
+(this is the end of the file)
+[File app.py created.]
+
+ASSISTANT:
+Now I will write the Python code for starting a web server and save it to the file `app.py`:
+<execute_ipython>
+EDITED_CODE=\"\"\"from flask import Flask
+app = Flask(__name__)
+
+@app.route('/')
+def index():
+    numbers = list(range(1, 11))
+    return str(numbers)
+
+if __name__ == '__main__':
+    app.run(port=5000)\"\"\"
+
+insert_content_at_line(
+  'app.py',
+  1,
+  EDITED_CODE,
+)
+</execute_ipython>
+
+USER:
+OBSERVATION:
+(this is the beginning of the file)
+1|from flask import Flask
+2|app = Flask(__name__)
+3|
+4|@app.route('/')
+5|def index():
+6|    numbers = list(range(1, 11))
+7|    return str(numbers)
+8|
+9|if __name__ == '__main__':
+10|    app.run(port=5000)
+(this is the end of the file)
+[File updated (edited at line 1). Please review the changes and make sure they are correct (correct indentation, no duplicate lines, etc). Edit the file again if necessary.]
+
+ASSISTANT:
+I have created a Python file `app.py` that will display a list of numbers from 1 to 10 when you run it. Let me run the Python file for you:
+<execute_bash>
+python3 app.py > server.log 2>&1 &
+</execute_bash>
+
+USER:
+OBSERVATION:
+[1] 121[1]+  Exit 1                  python3 app.py > server.log 2>&1
+
+ASSISTANT: Looks like the server is running with PID 121 then crashed. Let me check the server log:
+<execute_bash>
+cat server.log
+</execute_bash>
+
+USER:
+OBSERVATION:
+Traceback (most recent call last):
+  File "/workspace/app.py", line 2, in <module>
+    from flask import Flask
+ModuleNotFoundError: No module named 'flask'
+
+ASSISTANT:
+It seems that Flask is not installed. Let me install Flask for you:
+<execute_bash>
+pip install flask
+</execute_bash>
+
+USER:
+OBSERVATION:
+Defaulting to user installation because normal site-packages is not writeable
+Collecting flask
+  Using cached flask-3.0.3-py3-none-any.whl (101 kB)
+Collecting blinker>=1.6.2
+  Using cached blinker-1.7.0-py3-none-any.whl (13 kB)
+Collecting Werkzeug>=3.0.0
+  Using cached werkzeug-3.0.2-py3-none-any.whl (226 kB)
+Collecting click>=8.1.3
+  Using cached click-8.1.7-py3-none-any.whl (97 kB)
+Collecting itsdangerous>=2.1.2
+  Using cached itsdangerous-2.2.0-py3-none-any.whl (16 kB)
+Requirement already satisfied: Jinja2>=3.1.2 in /home/opendevin/.local/lib/python3.10/site-packages (from flask) (3.1.3)
+Requirement already satisfied: MarkupSafe>=2.0 in /home/opendevin/.local/lib/python3.10/site-packages (from Jinja2>=3.1.2->flask) (2.1.5)
+Installing collected packages: Werkzeug, itsdangerous, click, blinker, flask
+Successfully installed Werkzeug-3.0.2 blinker-1.7.0 click-8.1.7 flask-3.0.3 itsdangerous-2.2.0
+
+ASSISTANT:
+Now that Flask is installed, let me run the Python file again:
+<execute_bash>
+python3 app.py > server.log 2>&1 &
+</execute_bash>
+
+USER:
+OBSERVATION:
+[1] 124
+
+ASSISTANT:
+Let me check the server log again:
+<execute_bash>
+cat server.log
+</execute_bash>
+
+USER:
+OBSERVATION:
+* Serving Flask app 'app'
+ * Debug mode: off
+WARNING: This is a development server. Do not use it in a production deployment. Use a production WSGI server instead.
+ * Running on http://127.0.0.1:5000
+Press CTRL+C to quit
+
+ASSISTANT:
+The server is running on port 5000 with PID 124. You can access the list of numbers by visiting http://127.0.0.1:5000. If you have any further questions, feel free to ask!
+
+USER: Now browse the newly started server's homepage and show me the content.
+
+ASSISTANT:
+Sure! Let me browse the server's homepage at http://127.0.0.1:5000:
+<execute_browse>
+Get the content on "http://127.0.0.1:5000"
+</execute_browse>
+
+USER:
+OBSERVATION:
+[1, 2, 3, 4, 5, 6, 7, 8, 9, 10]
+
+ASSISTANT:
+The content of the server's homepage is "[1, 2, 3, 4, 5, 6, 7, 8, 9, 10]". If you have any further questions, feel free to ask!
+
+USER: Now kill the server, make it display the numbers in a table format.
+
+ASSISTANT:
+Sure! Let me stop the server first:
+<execute_bash>
+kill 124
+</execute_bash>
+
+USER:
+OBSERVATION:
+[1]+  Terminated              python3 app.py > server.log 2>&1
+
+ASSISTANT:
+The server has been stopped. Let me open the Python file and modify it.
+<execute_ipython>
+open_file('app.py')
+</execute_ipython>
+
+USER:
+[File: /workspace/app.py (10 lines total)]
+(this is the beginning of the file)
+1|from flask import Flask
+2|app = Flask(__name__)
+3|
+4|@app.route('/')
+5|def index():
+6|    numbers = list(range(1, 11))
+7|    return str(numbers)
+8|
+9|if __name__ == '__main__':
+10|    app.run(port=5000)
+(this is the end of the file)
+
+ASSISTANT:
+I should edit the file to display the numbers in a table format. I should include correct indentation. Let me update the file:
+<execute_ipython>
+edit_file_by_replace(
+  'app.py',
+  to_replace="    return str(numbers)",
+  new_content="    return '<table>' + ''.join([f'<tr><td>{i}</td></tr>' for i in numbers]) + '</table>'",
+)
+</execute_ipython>
+
+USER:
+Observation:
+[File: /workspace/app.py (10 lines total after edit)]
+(this is the beginning of the file)
+1|from flask import Flask
+2|app = Flask(__name__)
+3|
+4|@app.route('/')
+5|def index():
+6|    numbers = list(range(1, 11))
+7|    return '<table>' + ''.join([f'<tr><td>{i}</td></tr>' for i in numbers]) + '</table>'
+8|
+9|if __name__ == '__main__':
+10|    app.run(port=5000)
+(this is the end of the file)
+[File updated (edited at line 7). Please review the changes and make sure they are correct (correct indentation, no duplicate lines, etc). Edit the file again if necessary.]
+
+ASSISTANT:
+Running the updated file:
+<execute_bash>
+python3 app.py > server.log 2>&1 &
+</execute_bash>
+
+USER:
+Observation:
+[1] 126
+
+ASSISTANT:
+The server is running on port 5000 with PID 126. You can access the list of numbers in a table format by visiting http://127.0.0.1:5000. Let me know if you have any further requests!
+<finish></finish>
+
+--- END OF EXAMPLE ---
+"""
+
+INVALID_INPUT_MESSAGE = (
+    "I don't understand your input. \n"
+    'For bash commands, use <execute_bash> YOUR_COMMAND </execute_bash>.\n'
+    'For Python code, use <execute_ipython> YOUR_CODE </execute_ipython>.\n'
+    'For browsing, use <execute_browse> YOUR_COMMAND </execute_browse>.\n'
+)
--- a/agenthub/codeact_swe_agent/README.md
+++ b/agenthub/codeact_swe_agent/README.md
@@ -0,0 +1,7 @@
+# CodeAct (SWE Edit Specialized)
+
+This agent is an adaptation of the original [SWE Agent](https://swe-agent.com/) based on CodeAct using the `agentskills` library of OpenDevin.
+
+Its intended use is **solving GitHub issues**.
+
+It removes web-browsing and GitHub capability from the original CodeAct agent to avoid confusion to the agent.
--- a/agenthub/codeact_swe_agent/init.py
+++ b/agenthub/codeact_swe_agent/init.py
@@ -0,0 +1,5 @@
+from opendevin.controller.agent import Agent
+
+from .codeact_swe_agent import CodeActSWEAgent
+
+Agent.register('CodeActSWEAgent', CodeActSWEAgent)
--- a/agenthub/codeact_swe_agent/action_parser.py
+++ b/agenthub/codeact_swe_agent/action_parser.py
@@ -0,0 +1,114 @@
+import re
+
+from opendevin.controller.action_parser import ActionParser
+from opendevin.events.action import (
+    Action,
+    AgentFinishAction,
+    CmdRunAction,
+    IPythonRunCellAction,
+    MessageAction,
+)
+
+
+class CodeActSWEActionParserFinish(ActionParser):
+    """
+    Parser action:
+        - AgentFinishAction() - end the interaction
+    """
+
+    def __init__(
+        self,
+    ):
+        self.finish_command = None
+
+    def check_condition(self, action_str: str) -> bool:
+        self.finish_command = re.search(r'<finish>.*</finish>', action_str, re.DOTALL)
+        return self.finish_command is not None
+
+    def parse(self, action_str: str) -> Action:
+        assert (
+            self.finish_command is not None
+        ), 'self.finish_command should not be None when parse is called'
+        thought = action_str.replace(self.finish_command.group(0), '').strip()
+        return AgentFinishAction(thought=thought)
+
+
+class CodeActSWEActionParserCmdRun(ActionParser):
+    """
+    Parser action:
+        - CmdRunAction(command) - bash command to run
+        - AgentFinishAction() - end the interaction
+    """
+
+    def __init__(
+        self,
+    ):
+        self.bash_command = None
+
+    def check_condition(self, action_str: str) -> bool:
+        self.bash_command = re.search(
+            r'<execute_bash>(.*?)</execute_bash>', action_str, re.DOTALL
+        )
+        return self.bash_command is not None
+
+    def parse(self, action_str: str) -> Action:
+        assert (
+            self.bash_command is not None
+        ), 'self.bash_command should not be None when parse is called'
+        thought = action_str.replace(self.bash_command.group(0), '').strip()
+        # a command was found
+        command_group = self.bash_command.group(1).strip()
+        if command_group.strip() == 'exit':
+            return AgentFinishAction()
+        return CmdRunAction(command=command_group, thought=thought)
+
+
+class CodeActSWEActionParserIPythonRunCell(ActionParser):
+    """
+    Parser action:
+        - IPythonRunCellAction(code) - IPython code to run
+    """
+
+    def __init__(
+        self,
+    ):
+        self.python_code = None
+        self.jupyter_kernel_init_code: str = 'from agentskills import *'
+
+    def check_condition(self, action_str: str) -> bool:
+        self.python_code = re.search(
+            r'<execute_ipython>(.*?)</execute_ipython>', action_str, re.DOTALL
+        )
+        return self.python_code is not None
+
+    def parse(self, action_str: str) -> Action:
+        assert (
+            self.python_code is not None
+        ), 'self.python_code should not be None when parse is called'
+        code_group = self.python_code.group(1).strip()
+        thought = action_str.replace(self.python_code.group(0), '').strip()
+        return IPythonRunCellAction(
+            code=code_group,
+            thought=thought,
+            kernel_init_code=self.jupyter_kernel_init_code,
+        )
+
+
+class CodeActSWEActionParserMessage(ActionParser):
+    """
+    Parser action:
+        - MessageAction(content) - Message action to run (e.g. ask for clarification)
+    """
+
+    def __init__(
+        self,
+    ):
+        pass
+
+    def check_condition(self, action_str: str) -> bool:
+        # We assume the LLM is GOOD enough that when it returns pure natural language
+        # it wants to talk to the user
+        return True
+
+    def parse(self, action_str: str) -> Action:
+        return MessageAction(content=action_str, wait_for_response=True)
--- a/agenthub/codeact_swe_agent/codeact_swe_agent.py
+++ b/agenthub/codeact_swe_agent/codeact_swe_agent.py
@@ -0,0 +1,195 @@
+from agenthub.codeact_swe_agent.prompt import (
+    COMMAND_DOCS,
+    SWE_EXAMPLE,
+    SYSTEM_PREFIX,
+    SYSTEM_SUFFIX,
+)
+from agenthub.codeact_swe_agent.response_parser import CodeActSWEResponseParser
+from opendevin.controller.agent import Agent
+from opendevin.controller.state.state import State
+from opendevin.core.config import config
+from opendevin.events.action import (
+    Action,
+    AgentFinishAction,
+    CmdRunAction,
+    IPythonRunCellAction,
+    MessageAction,
+)
+from opendevin.events.observation import (
+    CmdOutputObservation,
+    IPythonRunCellObservation,
+)
+from opendevin.events.serialization.event import truncate_content
+from opendevin.llm.llm import LLM
+from opendevin.runtime.plugins import (
+    AgentSkillsRequirement,
+    JupyterRequirement,
+    PluginRequirement,
+)
+from opendevin.runtime.tools import RuntimeTool
+
+
+def action_to_str(action: Action) -> str:
+    if isinstance(action, CmdRunAction):
+        return f'{action.thought}\n<execute_bash>\n{action.command}\n</execute_bash>'
+    elif isinstance(action, IPythonRunCellAction):
+        return f'{action.thought}\n<execute_ipython>\n{action.code}\n</execute_ipython>'
+    elif isinstance(action, MessageAction):
+        return action.content
+    return ''
+
+
+def get_action_message(action: Action) -> dict[str, str] | None:
+    if (
+        isinstance(action, CmdRunAction)
+        or isinstance(action, IPythonRunCellAction)
+        or isinstance(action, MessageAction)
+    ):
+        return {
+            'role': 'user' if action.source == 'user' else 'assistant',
+            'content': action_to_str(action),
+        }
+    return None
+
+
+def get_observation_message(obs) -> dict[str, str] | None:
+    max_message_chars = config.get_llm_config_from_agent(
+        'CodeActSWEAgent'
+    ).max_message_chars
+    if isinstance(obs, CmdOutputObservation):
+        content = 'OBSERVATION:\n' + truncate_content(obs.content, max_message_chars)
+        content += (
+            f'\n[Command {obs.command_id} finished with exit code {obs.exit_code}]'
+        )
+        return {'role': 'user', 'content': content}
+    elif isinstance(obs, IPythonRunCellObservation):
+        content = 'OBSERVATION:\n' + obs.content
+        # replace base64 images with a placeholder
+        splitted = content.split('\n')
+        for i, line in enumerate(splitted):
+            if '![image](data:image/png;base64,' in line:
+                splitted[i] = (
+                    '![image](data:image/png;base64, ...) already displayed to user'
+                )
+        content = '\n'.join(splitted)
+        content = truncate_content(content, max_message_chars)
+        return {'role': 'user', 'content': content}
+    return None
+
+
+def get_system_message() -> str:
+    return f'{SYSTEM_PREFIX}\n\n{COMMAND_DOCS}\n\n{SYSTEM_SUFFIX}'
+
+
+def get_in_context_example() -> str:
+    return SWE_EXAMPLE
+
+
+class CodeActSWEAgent(Agent):
+    VERSION = '1.6'
+    """
+    This agent is an adaptation of the original [SWE Agent](https://swe-agent.com/) based on CodeAct 1.5 using the `agentskills` library of OpenDevin.
+
+    It is intended use is **solving Github issues**.
+
+    It removes web-browsing and Github capability from the original CodeAct agent to avoid confusion to the agent.
+    """
+
+    sandbox_plugins: list[PluginRequirement] = [
+        # NOTE: AgentSkillsRequirement need to go before JupyterRequirement, since
+        # AgentSkillsRequirement provides a lot of Python functions,
+        # and it needs to be initialized before Jupyter for Jupyter to use those functions.
+        AgentSkillsRequirement(),
+        JupyterRequirement(),
+    ]
+    runtime_tools: list[RuntimeTool] = []
+
+    system_message: str = get_system_message()
+    in_context_example: str = f"Here is an example of how you can interact with the environment for task solving:\n{get_in_context_example()}\n\nNOW, LET'S START!"
+
+    response_parser = CodeActSWEResponseParser()
+
+    def __init__(
+        self,
+        llm: LLM,
+    ) -> None:
+        """
+        Initializes a new instance of the CodeActAgent class.
+
+        Parameters:
+        - llm (LLM): The llm to be used by this agent
+        """
+        super().__init__(llm)
+        self.reset()
+
+    def reset(self) -> None:
+        """
+        Resets the CodeAct Agent.
+        """
+        super().reset()
+
+    def step(self, state: State) -> Action:
+        """
+        Performs one step using the CodeAct Agent.
+        This includes gathering info on previous steps and prompting the model to make a command to execute.
+
+        Parameters:
+        - state (State): used to get updated info and background commands
+
+        Returns:
+        - CmdRunAction(command) - bash command to run
+        - IPythonRunCellAction(code) - IPython code to run
+        - MessageAction(content) - Message action to run (e.g. ask for clarification)
+        - AgentFinishAction() - end the interaction
+        """
+
+        # if we're done, go back
+        latest_user_message = state.history.get_last_user_message()
+        if latest_user_message and latest_user_message.strip() == '/exit':
+            return AgentFinishAction()
+
+        # prepare what we want to send to the LLM
+        messages: list[dict[str, str]] = self._get_messages(state)
+
+        response = self.llm.completion(
+            messages=messages,
+            stop=[
+                '</execute_ipython>',
+                '</execute_bash>',
+            ],
+            temperature=0.0,
+        )
+
+        return self.response_parser.parse(response)
+
+    def _get_messages(self, state: State) -> list[dict[str, str]]:
+        messages = [
+            {'role': 'system', 'content': self.system_message},
+            {'role': 'user', 'content': self.in_context_example},
+        ]
+
+        for event in state.history.get_events():
+            # create a regular message from an event
+            message = (
+                get_action_message(event)
+                if isinstance(event, Action)
+                else get_observation_message(event)
+            )
+
+            # add regular message
+            if message:
+                messages.append(message)
+
+        # the latest user message is important:
+        # we want to remind the agent of the environment constraints
+        latest_user_message = next(
+            (m for m in reversed(messages) if m['role'] == 'user'), None
+        )
+
+        # add a reminder to the prompt
+        if latest_user_message:
+            latest_user_message['content'] += (
+                f'\n\nENVIRONMENT REMINDER: You have {state.max_iterations - state.iteration} turns left to complete the task.'
+            )
+
+        return messages
--- a/agenthub/codeact_swe_agent/prompt.py
+++ b/agenthub/codeact_swe_agent/prompt.py
@@ -0,0 +1,455 @@
+from opendevin.runtime.plugins import AgentSkillsRequirement
+
+_AGENT_SKILLS_DOCS = AgentSkillsRequirement.documentation
+
+COMMAND_DOCS = (
+    '\nApart from the standard Python library, the assistant can also use the following functions (already imported) in <execute_ipython> environment:\n'
+    f'{_AGENT_SKILLS_DOCS}'
+    "Please note that THE `edit_file` FUNCTION REQUIRES PROPER INDENTATION. If the assistant would like to add the line '        print(x)', it must fully write that out, with all those spaces before the code! Indentation is important and code that is not indented correctly will fail and require fixing before it can be run."
+)
+
+# ======= SYSTEM MESSAGE =======
+MINIMAL_SYSTEM_PREFIX = """A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions.
+The assistant can interact with an interactive Python (Jupyter Notebook) environment and receive the corresponding output when needed. The code should be enclosed using "<execute_ipython>" tag, for example:
+<execute_ipython>
+print("Hello World!")
+</execute_ipython>
+The assistant can execute bash commands on behalf of the user by wrapping them with <execute_bash> and </execute_bash>.
+For example, you can list the files in the current directory by <execute_bash> ls </execute_bash>.
+"""
+
+PIP_INSTALL_PREFIX = """The assistant can install Python packages using the %pip magic command in an IPython environment by using the following syntax: <execute_ipython> %pip install [package needed] </execute_ipython> and should always import packages and define variables before starting to use them."""
+
+SYSTEM_PREFIX = MINIMAL_SYSTEM_PREFIX + PIP_INSTALL_PREFIX
+
+SYSTEM_SUFFIX = """The assistant's response should be concise.
+The assistant should include ONLY ONE <execute_ipython> or <execute_bash> in every one of the responses, unless the assistant is finished with the task or need more input or action from the user in order to proceed.
+IMPORTANT: Whenever possible, execute the code for the user using <execute_ipython> or <execute_bash> instead of providing it.
+"""
+
+SWE_EXAMPLE = """
+--- START OF EXAMPLE ---
+
+USER: We are currently solving the following issue within our repository. Here is the issue text:
+--- BEGIN ISSUE ---
+TimeDelta serialization precision
+Hi there!
+
+I just found quite strange behaviour of `TimeDelta` field serialization
+
+```python3
+from marshmallow.fields import TimeDelta
+from datetime import timedelta
+
+td_field = TimeDelta(precision="milliseconds")
+
+obj = dict()
+obj["td_field"] = timedelta(milliseconds=345)
+
+print(td_field.serialize("td_field", obj))
+```
+
+Output of this snippet is `344`, but it seems that `345` is correct.
+
+Looks like a rounding issue here: https://github.com/marshmallow-code/marshmallow/blob/dev/src/marshmallow/fields.py#L1474
+--- END ISSUE ---
+
+[Current directory: /workspace/marshmallow-code__marshmallow]
+
+ASSISTANT:
+Let's first start by reproducing the results of the issue. The issue includes some example code for reproduction, which we can use. We'll create a new file called `reproduce.py` and paste the example code into it.
+<execute_ipython>
+create_file('/workspace/marshmallow-code__marshmallow/reproduce.py')
+</execute_ipython>
+
+USER:
+OBSERVATION:
+[File: /workspace/marshmallow-code__marshmallow/reproduce.py (1 lines total)]
+1|
+[File /workspace/marshmallow-code__marshmallow/reproduce.py created.]
+
+ASSISTANT:
+Now let's paste in the example code from the issue.
+<execute_ipython>
+EDITED_CODE = \"\"\"from marshmallow.fields import TimeDelta
+from datetime import timedelta
+
+td_field = TimeDelta(precision="milliseconds")
+
+obj = dict()
+obj["td_field"] = timedelta(milliseconds=345)
+
+print(td_field.serialize("td_field", obj))
+\"\"\"
+edit_file(start=1, end=1, content=EDITED_CODE)
+</execute_ipython>
+
+USER:
+OBSERVATION:
+[File: /workspace/marshmallow-code__marshmallow/reproduce.py (9 lines total)]
+1|from marshmallow.fields import TimeDelta
+2|from datetime import timedelta
+3|
+4|td_field = TimeDelta(precision="milliseconds")
+5|
+6|obj = dict()
+7|obj["td_field"] = timedelta(milliseconds=345)
+8|
+9|print(td_field.serialize("td_field", obj))
+[File updated. Please review the changes and make sure they are correct (correct indentation, no duplicate lines, etc). Edit the file again if necessary.]
+
+ASSISTANT:
+Now let's run the code to see if we see the same output as the issue.
+<execute_bash>
+python reproduce.py
+</execute_bash>
+
+USER:
+OBSERVATION:
+344
+
+ASSISTANT:
+We are indeed seeing the same output as the issue. The issue suggests that we should look at line 1474 of the `fields.py` file to see if there is a rounding issue. Let's open the `fields.py` file and navigate to line 1474 to see the relevant code. First we should find the `fields.py` file in the repository. It is likely to be in the `src/marshmallow` directory, but we should check to be sure.
+<execute_bash>
+ls -F
+</execute_bash>
+
+USER:
+OBSERVATION:
+AUTHORS.rst
+CHANGELOG.rst
+CODE_OF_CONDUCT.md
+CONTRIBUTING.rst
+LICENSE
+MANIFEST.in
+NOTICE
+README.rst
+RELEASING.md
+azure-pipelines.yml
+docs/
+examples/
+performance/
+pyproject.toml
+reproduce.py
+setup.cfg
+setup.py
+src/
+tests/
+tox.ini
+
+ASSISTANT:
+It looks like the `src` directory is present, which suggests that the `fields.py` file is likely to be in the `src` directory. Let's use find_file to see where it is.
+<execute_ipython>
+find_file("fields.py", "/workspace/marshmallow-code__marshmallow/src")
+</execute_ipython>
+
+USER:
+OBSERVATION:
+[Found 1 matches for "fields.py" in /workspace/marshmallow-code__marshmallow/src]
+/workspace/marshmallow-code__marshmallow/src/marshmallow/fields.py
+
+ASSISTANT:
+It looks like the `fields.py` file is present in the `./src/marshmallow/` directory. The issue also points to a specific URL with line number 1474. We should navigate to that line in fields.py to see the relevant code for the `TimeDelta` serialization.
+<execute_ipython>
+open_file("/workspace/marshmallow-code__marshmallow/src/marshmallow/fields.py", 1474)
+</execute_ipython>
+
+USER:
+OBSERVATION:
+[File: /workspace/marshmallow-code__marshmallow/src/marshmallow/fields.py (1997 lines total)]
+(1458 more lines above)
+1459|            self.WEEKS,
+1460|        )
+1461|
+1462|        if precision not in units:
+1463|            msg = 'The precision must be {} or "{}".'.format(
+1464|                ", ".join([f'"{each}"' for each in units[:-1]]), units[-1]
+1465|            )
+1466|            raise ValueError(msg)
+1467|
+1468|        self.precision = precision
+1469|        super().__init__(**kwargs)
+1470|
+1471|    def _serialize(self, value, attr, obj, **kwargs):
+1472|        if value is None:
+1473|            return None
+1474|        base_unit = dt.timedelta(**{self.precision: 1})
+1475|        return int(value.total_seconds() / base_unit.total_seconds())
+1476|
+1477|    def _deserialize(self, value, attr, data, **kwargs):
+1478|        try:
+1479|            value = int(value)
+1480|        except (TypeError, ValueError) as error:
+1481|            raise self.make_error("invalid") from error
+1482|
+1483|        kwargs = {self.precision: value}
+1484|
+1485|        try:
+1486|            return dt.timedelta(**kwargs)
+1487|        except OverflowError as error:
+1488|            raise self.make_error("invalid") from error
+1489|
+1490|
+1491|class Mapping(Field):
+1492|    \"\"\"An abstract class for objects with key-value pairs.
+1493|
+1494|    :param keys: A field class or instance for dict keys.
+1495|    :param values: A field class or instance for dict values.
+1496|    :param kwargs: The same keyword arguments that :class:`Field` receives.
+1497|
+1498|    .. note::
+1499|        When the structure of nested data is not known, you may omit the
+1500|        `keys` and `values` arguments to prevent content validation.
+1501|
+1502|    .. versionadded:: 3.0.0rc4
+1503|    \"\"\"
+1504|
+1505|    mapping_type = dict
+1506|
+1507|    #: Default error messages.
+1508|    default_error_messages = {"invalid": "Not a valid mapping type."}
+1509|
+1510|    def __init__(
+1511|        self,
+1512|        keys: typing.Optional[typing.Union[Field, type]] = None,
+1513|        values: typing.Optional[typing.Union[Field, type]] = None,
+1514|        **kwargs,
+1515|    ):
+1516|        super().__init__(**kwargs)
+1517|        if keys is None:
+1518|            self.key_field = None
+1519|        else:
+1520|            try:
+1521|                self.key_field = resolve_field_instance(keys)
+1522|            except FieldInstanceResolutionError as error:
+1523|                raise ValueError(
+1524|                    '"keys" must be a subclass or instance of '
+1525|                    "marshmallow.base.FieldABC."
+1526|                ) from error
+1527|
+1528|        if values is None:
+1529|            self.value_field = None
+1530|        else:
+1531|            try:
+1532|                self.value_field = resolve_field_instance(values)
+1533|            except FieldInstanceResolutionError as error:
+1534|                raise ValueError(
+1535|                    '"values" must be a subclass or instance of '
+1536|                    "marshmallow.base.FieldABC."
+1537|                ) from error
+1538|            if isinstance(self.value_field, Nested):
+1539|                self.only = self.value_field.only
+1540|                self.exclude = self.value_field.exclude
+1541|
+1542|    def _bind_to_schema(self, field_name, schema):
+1543|        super()._bind_to_schema(field_name, schema)
+1544|        if self.value_field:
+1545|            self.value_field = copy.deepcopy(self.value_field)
+1546|            self.value_field._bind_to_schema(field_name, self)
+1547|        if isinstance(self.value_field, Nested):
+1548|            self.value_field.only = self.only
+1549|            self.value_field.exclude = self.exclude
+1550|        if self.key_field:
+1551|            self.key_field = copy.deepcopy(self.key_field)
+1552|            self.key_field._bind_to_schema(field_name, self)
+1553|
+1554|    def _serialize(self, value, attr, obj, **kwargs):
+1555|        if value is None:
+1556|            return None
+1557|        if not self.value_field and not self.key_field:
+1558|            return self.mapping_type(value)
+(439 more lines below)
+
+ASSISTANT:
+We are now looking at the relevant section of the `fields.py` file where the `TimeDelta` serialization occurs. The issue suggests that there is a rounding problem near line 1474. The current implementation uses integer division, which truncates the decimal part and can lead to incorrect rounding on line 1475. To fix this issue, we should ensure that the division results in a float, and then round the result to the nearest integer before converting it to an int. This should preserve the milliseconds precision as expected. Let's make the necessary edit to the code.
+<execute_ipython>
+edit_file(start=1475, end=1475, content="return int(round(value.total_seconds() / base_unit.total_seconds()))  # round to nearest int")
+</execute_ipython>
+
+USER:
+OBSERVATION:
+[Your proposed edit has introduced new syntax error(s). Please understand the errors and retry your edit command.]
+
+ERRORS:
+- E999 IndentationError: unexpected indent
+
+[This is how your edit would have looked if applied]
+-------------------------------------------------
+[File: /marshmallow-code__marshmallow/src/marshmallow/fields.py (1997 lines total)]
+(1468 more lines above)
+1469|        super().__init__(**kwargs)
+1470|
+1471|    def _serialize(self, value, attr, obj, **kwargs):
+1472|        if value is None:
+1473|            return None
+1474|        base_unit = dt.timedelta(**{self.precision: 1})
+1475|return int(round(value.total_seconds() / base_unit.total_seconds()))  # round to nearest int
+1476|
+1477|    def _deserialize(self, value, attr, data, **kwargs):
+1478|        try:
+1479|            value = int(value)
+(518 more lines below)
+-------------------------------------------------
+
+[This is the original code before your edit]
+-------------------------------------------------
+[File: /marshmallow-code__marshmallow/src/marshmallow/fields.py (1997 lines total)]
+(1469 more lines above)
+1470|
+1471|    def _serialize(self, value, attr, obj, **kwargs):
+1472|        if value is None:
+1473|            return None
+1474|        base_unit = dt.timedelta(**{self.precision: 1})
+1475|        return int(value.total_seconds() / base_unit.total_seconds())
+1476|
+1477|    def _deserialize(self, value, attr, data, **kwargs):
+1478|        try:
+1479|            value = int(value)
+1480|        except (TypeError, ValueError) as error:
+(517 more lines below)
+-------------------------------------------------
+Your changes have NOT been applied. Please fix your edit command and try again.
+You either need to 1) Specify the correct start/end line arguments or 2) Correct your edit code.
+DO NOT re-run the same failed edit command. Running it again will lead to the same error.
+
+ASSISTANT:
+My edit command did not use the proper indentation, I will fix my syntax in this follow up edit command.
+<execute_ipython>
+EDITED_CODE="        return int(round(value.total_seconds() / base_unit.total_seconds()))  # round to nearest int"
+edit_file(start=1475, end=1475, content=EDITED_CODE)
+</execute_ipython>
+
+USER:
+OBSERVATION:
+[File: /marshmallow-code__marshmallow/src/marshmallow/fields.py (1997 lines total)]
+(1424 more lines above)
+1425|
+1426|    :param precision: Influences how the integer is interpreted during
+1427|        (de)serialization. Must be 'days', 'seconds', 'microseconds',
+1428|        'milliseconds', 'minutes', 'hours' or 'weeks'.
+1429|    :param kwargs: The same keyword arguments that :class:`Field` receives.
+1430|
+1431|    .. versionchanged:: 2.0.0
+1432|        Always serializes to an integer value to avoid rounding errors.
+1433|        Add `precision` parameter.
+1434|    \"\"\"
+1435|
+1436|    DAYS = "days"
+1437|    SECONDS = "seconds"
+1438|    MICROSECONDS = "microseconds"
+1439|    MILLISECONDS = "milliseconds"
+1440|    MINUTES = "minutes"
+1441|    HOURS = "hours"
+1442|    WEEKS = "weeks"
+1443|
+1444|    #: Default error messages.
+1445|    default_error_messages = {
+1446|        "invalid": "Not a valid period of time.",
+1447|        "format": "{input!r} cannot be formatted as a timedelta.",
+1448|    }
+1449|
+1450|    def __init__(self, precision: str = SECONDS, **kwargs):
+1451|        precision = precision.lower()
+1452|        units = (
+1453|            self.DAYS,
+1454|            self.SECONDS,
+1455|            self.MICROSECONDS,
+1456|            self.MILLISECONDS,
+1457|            self.MINUTES,
+1458|            self.HOURS,
+1459|            self.WEEKS,
+1460|        )
+1461|
+1462|        if precision not in units:
+1463|            msg = 'The precision must be {} or "{}".'.format(
+1464|                ", ".join([f'"{each}"' for each in units[:-1]]), units[-1]
+1465|            )
+1466|            raise ValueError(msg)
+1467|
+1468|        self.precision = precision
+1469|        super().__init__(**kwargs)
+1470|
+1471|    def _serialize(self, value, attr, obj, **kwargs):
+1472|        if value is None:
+1473|            return None
+1474|        base_unit = dt.timedelta(**{self.precision: 1})
+1475|        return int(round(value.total_seconds() / base_unit.total_seconds()))  # round to nearest int
+1476|
+1477|    def _deserialize(self, value, attr, data, **kwargs):
+1478|        try:
+1479|            value = int(value)
+1480|        except (TypeError, ValueError) as error:
+1481|            raise self.make_error("invalid") from error
+1482|
+1483|        kwargs = {self.precision: value}
+1484|
+1485|        try:
+1486|            return dt.timedelta(**kwargs)
+1487|        except OverflowError as error:
+1488|            raise self.make_error("invalid") from error
+1489|
+1490|
+1491|class Mapping(Field):
+1492|    \"\"\"An abstract class for objects with key-value pairs.
+1493|
+1494|    :param keys: A field class or instance for dict keys.
+1495|    :param values: A field class or instance for dict values.
+1496|    :param kwargs: The same keyword arguments that :class:`Field` receives.
+1497|
+1498|    .. note::
+1499|        When the structure of nested data is not known, you may omit the
+1500|        `keys` and `values` arguments to prevent content validation.
+1501|
+1502|    .. versionadded:: 3.0.0rc4
+1503|    \"\"\"
+1504|
+1505|    mapping_type = dict
+1506|
+1507|    #: Default error messages.
+1508|    default_error_messages = {"invalid": "Not a valid mapping type."}
+1509|
+1510|    def __init__(
+1511|        self,
+1512|        keys: typing.Optional[typing.Union[Field, type]] = None,
+1513|        values: typing.Optional[typing.Union[Field, type]] = None,
+1514|        **kwargs,
+1515|    ):
+1516|        super().__init__(**kwargs)
+1517|        if keys is None:
+1518|            self.key_field = None
+1519|        else:
+1520|            try:
+1521|                self.key_field = resolve_field_instance(keys)
+1522|            except FieldInstanceResolutionError as error:
+1523|                raise ValueError(
+1524|                    '"keys" must be a subclass or instance of '
+(473 more lines below)
+[File updated. Please review the changes and make sure they are correct (correct indentation, no duplicate lines, etc). Edit the file again if necessary.]
+
+ASSISTANT:
+The code has been updated to use the `round` function, which should fix the rounding issue. Before submitting the changes, it would be prudent to run the reproduce.py code again to ensure that this change has the desired effect. I see that "Current directory" is still `/marshmallow-code__marshmallow`, so we can just run `reproduce.py` directly.
+<execute_bash>
+python reproduce.py
+</execute_bash>
+
+USER:
+OBSERVATION:
+345
+
+ASSISTANT:
+The output has changed from 344 to 345, which suggests that the rounding issue has been fixed. Let's remove the reproduce.py file since it is no longer needed.
+<execute_bash>
+rm reproduce.py
+</execute_bash>
+
+USER:
+OBSERVATION:
+
+ASSISTANT:
+rm doesn't have any output when it deletes successfully, so that must have worked. Now that we have fixed the issue, we can exit the current task.
+<execute_bash>
+exit
+</execute_bash>
+
+--- END OF EXAMPLE ---
+"""
--- a/agenthub/codeact_swe_agent/response_parser.py
+++ b/agenthub/codeact_swe_agent/response_parser.py
@@ -0,0 +1,47 @@
+from agenthub.codeact_swe_agent.action_parser import (
+    CodeActSWEActionParserCmdRun,
+    CodeActSWEActionParserFinish,
+    CodeActSWEActionParserIPythonRunCell,
+    CodeActSWEActionParserMessage,
+)
+from opendevin.controller.action_parser import ResponseParser
+from opendevin.events.action import Action
+
+
+class CodeActSWEResponseParser(ResponseParser):
+    """
+    Parser action:
+        - CmdRunAction(command) - bash command to run
+        - IPythonRunCellAction(code) - IPython code to run
+        - MessageAction(content) - Message action to run (e.g. ask for clarification)
+        - AgentFinishAction() - end the interaction
+    """
+
+    def __init__(self):
+        # Need pay attention to the item order in self.action_parsers
+        super().__init__()
+        self.action_parsers = [
+            CodeActSWEActionParserFinish(),
+            CodeActSWEActionParserCmdRun(),
+            CodeActSWEActionParserIPythonRunCell(),
+        ]
+        self.default_parser = CodeActSWEActionParserMessage()
+
+    def parse(self, response: str) -> Action:
+        action_str = self.parse_response(response)
+        return self.parse_action(action_str)
+
+    def parse_response(self, response) -> str:
+        action = response.choices[0].message.content
+        if action is None:
+            return ''
+        for lang in ['bash', 'ipython']:
+            if f'<execute_{lang}>' in action and f'</execute_{lang}>' not in action:
+                action += f'</execute_{lang}>'
+        return action
+
+    def parse_action(self, action_str: str) -> Action:
+        for action_parser in self.action_parsers:
+            if action_parser.check_condition(action_str):
+                return action_parser.parse(action_str)
+        return self.default_parser.parse(action_str)
--- a/agenthub/delegator_agent/init.py
+++ b/agenthub/delegator_agent/init.py
@@ -0,0 +1,5 @@
+from opendevin.controller.agent import Agent
+
+from .agent import DelegatorAgent
+
+Agent.register('DelegatorAgent', DelegatorAgent)
--- a/agenthub/delegator_agent/agent.py
+++ b/agenthub/delegator_agent/agent.py
@@ -0,0 +1,84 @@
+from opendevin.controller.agent import Agent
+from opendevin.controller.state.state import State
+from opendevin.events.action import Action, AgentDelegateAction, AgentFinishAction
+from opendevin.events.observation import AgentDelegateObservation
+from opendevin.llm.llm import LLM
+
+
+class DelegatorAgent(Agent):
+    VERSION = '1.0'
+    """
+    The Delegator Agent is responsible for delegating tasks to other agents based on the current task.
+    """
+
+    current_delegate: str = ''
+
+    def __init__(self, llm: LLM):
+        """
+        Initialize the Delegator Agent with an LLM
+
+        Parameters:
+        - llm (LLM): The llm to be used by this agent
+        """
+        super().__init__(llm)
+
+    def step(self, state: State) -> Action:
+        """
+        Checks to see if current step is completed, returns AgentFinishAction if True.
+        Otherwise, delegates the task to the next agent in the pipeline.
+
+        Parameters:
+        - state (State): The current state given the previous actions and observations
+
+        Returns:
+        - AgentFinishAction: If the last state was 'completed', 'verified', or 'abandoned'
+        - AgentDelegateAction: The next agent to delegate the task to
+        """
+        if self.current_delegate == '':
+            self.current_delegate = 'study'
+            task = state.get_current_user_intent()
+            return AgentDelegateAction(
+                agent='StudyRepoForTaskAgent', inputs={'task': task}
+            )
+
+        # last observation in history should be from the delegate
+        last_observation = state.history.get_last_observation()
+
+        if not isinstance(last_observation, AgentDelegateObservation):
+            raise Exception('Last observation is not an AgentDelegateObservation')
+
+        goal = state.get_current_user_intent()
+        if self.current_delegate == 'study':
+            self.current_delegate = 'coder'
+            return AgentDelegateAction(
+                agent='CoderAgent',
+                inputs={
+                    'task': goal,
+                    'summary': last_observation.outputs['summary'],
+                },
+            )
+        elif self.current_delegate == 'coder':
+            self.current_delegate = 'verifier'
+            return AgentDelegateAction(
+                agent='VerifierAgent',
+                inputs={
+                    'task': goal,
+                },
+            )
+        elif self.current_delegate == 'verifier':
+            if (
+                'completed' in last_observation.outputs
+                and last_observation.outputs['completed']
+            ):
+                return AgentFinishAction()
+            else:
+                self.current_delegate = 'coder'
+                return AgentDelegateAction(
+                    agent='CoderAgent',
+                    inputs={
+                        'task': goal,
+                        'summary': last_observation.outputs['summary'],
+                    },
+                )
+        else:
+            raise Exception('Invalid delegate state')
--- a/agenthub/dummy_agent/init.py
+++ b/agenthub/dummy_agent/init.py
@@ -0,0 +1,5 @@
+from opendevin.controller.agent import Agent
+
+from .agent import DummyAgent
+
+Agent.register('DummyAgent', DummyAgent)
--- a/agenthub/dummy_agent/agent.py
+++ b/agenthub/dummy_agent/agent.py
@@ -0,0 +1,146 @@
+import time
+from typing import TypedDict
+
+from opendevin.controller.agent import Agent
+from opendevin.controller.state.state import State
+from opendevin.events.action import (
+    Action,
+    AddTaskAction,
+    AgentFinishAction,
+    AgentRejectAction,
+    BrowseInteractiveAction,
+    BrowseURLAction,
+    CmdRunAction,
+    FileReadAction,
+    FileWriteAction,
+    MessageAction,
+    ModifyTaskAction,
+)
+from opendevin.events.observation import (
+    CmdOutputObservation,
+    FileReadObservation,
+    FileWriteObservation,
+    NullObservation,
+    Observation,
+)
+from opendevin.events.serialization.event import event_to_dict
+from opendevin.llm.llm import LLM
+
+"""
+FIXME: There are a few problems this surfaced
+* FileWrites seem to add an unintended newline at the end of the file
+* Browser not working
+"""
+
+ActionObs = TypedDict(
+    'ActionObs', {'action': Action, 'observations': list[Observation]}
+)
+
+
+class DummyAgent(Agent):
+    VERSION = '1.0'
+    """
+    The DummyAgent is used for e2e testing. It just sends the same set of actions deterministically,
+    without making any LLM calls.
+    """
+
+    def __init__(self, llm: LLM):
+        super().__init__(llm)
+        self.steps: list[ActionObs] = [
+            {
+                'action': AddTaskAction(parent='0', goal='check the current directory'),
+                'observations': [NullObservation('')],
+            },
+            {
+                'action': AddTaskAction(parent='0.0', goal='run ls'),
+                'observations': [NullObservation('')],
+            },
+            {
+                'action': ModifyTaskAction(task_id='0.0', state='in_progress'),
+                'observations': [NullObservation('')],
+            },
+            {
+                'action': MessageAction('Time to get started!'),
+                'observations': [NullObservation('')],
+            },
+            {
+                'action': CmdRunAction(command='echo "foo"'),
+                'observations': [
+                    CmdOutputObservation('foo', command_id=-1, command='echo "foo"')
+                ],
+            },
+            {
+                'action': FileWriteAction(
+                    content='echo "Hello, World!"', path='hello.sh'
+                ),
+                'observations': [FileWriteObservation('', path='hello.sh')],
+            },
+            {
+                'action': FileReadAction(path='hello.sh'),
+                'observations': [
+                    FileReadObservation('echo "Hello, World!"\n', path='hello.sh')
+                ],
+            },
+            {
+                'action': CmdRunAction(command='bash hello.sh'),
+                'observations': [
+                    CmdOutputObservation(
+                        'Hello, World!', command_id=-1, command='bash hello.sh'
+                    )
+                ],
+            },
+            {
+                'action': BrowseURLAction(url='https://google.com'),
+                'observations': [
+                    # BrowserOutputObservation('<html></html>', url='https://google.com', screenshot=""),
+                ],
+            },
+            {
+                'action': BrowseInteractiveAction(
+                    browser_actions='goto("https://google.com")'
+                ),
+                'observations': [
+                    # BrowserOutputObservation('<html></html>', url='https://google.com', screenshot=""),
+                ],
+            },
+            {
+                'action': AgentFinishAction(),
+                'observations': [],
+            },
+            {
+                'action': AgentRejectAction(),
+                'observations': [],
+            },
+        ]
+
+    def step(self, state: State) -> Action:
+        time.sleep(0.1)
+        if state.iteration > 0:
+            prev_step = self.steps[state.iteration - 1]
+
+            # a step is (action, observations list)
+            if 'observations' in prev_step:
+                # one obs, at most
+                expected_observations = prev_step['observations']
+
+                # check if the history matches the expected observations
+                hist_events = state.history.get_last_events(len(expected_observations))
+                for i in range(len(expected_observations)):
+                    hist_obs = event_to_dict(hist_events[i])
+                    expected_obs = event_to_dict(expected_observations[i])
+                    if (
+                        'command_id' in hist_obs['extras']
+                        and hist_obs['extras']['command_id'] != -1
+                    ):
+                        del hist_obs['extras']['command_id']
+                        hist_obs['content'] = ''
+                    if (
+                        'command_id' in expected_obs['extras']
+                        and expected_obs['extras']['command_id'] != -1
+                    ):
+                        del expected_obs['extras']['command_id']
+                        expected_obs['content'] = ''
+                    assert (
+                        hist_obs == expected_obs
+                    ), f'Expected observation {expected_obs}, got {hist_obs}'
+        return self.steps[state.iteration]['action']
--- a/agenthub/gptswarm_agent/README.md
+++ b/agenthub/gptswarm_agent/README.md
@@ -0,0 +1,16 @@
+# GPTSwarm Framework
+
+## Introduction
+
+This folder implements the GPTSwarm ([paper](https://arxiv.org/abs/2402.01030), [Original Repo](https://github.com/metauto-ai/GPTSwarm)).  For more details, please see paper.
+
+
+## Reference
+```
+@article{zhuge2024language,
+  title={Language Agents as Optimizable Graphs},
+  author={Zhuge, Mingchen and Wang, Wenyi and Kirsch, Louis and Faccio, Francesco and Khizbullin, Dmitrii and Schmidhuber, Jurgen},
+  journal={arXiv preprint arXiv:2402.16823},
+  year={2024}
+}
+```
--- a/agenthub/gptswarm_agent/init.py
+++ b/agenthub/gptswarm_agent/init.py
@@ -0,0 +1,5 @@
+from opendevin.controller.agent import Agent
+
+from .gptswarm_agent import GPTSwarm
+
+Agent.register('GPTSwarmAgent', GPTSwarm)
--- a/agenthub/gptswarm_agent/gptswarm_agent.py
+++ b/agenthub/gptswarm_agent/gptswarm_agent.py
@@ -0,0 +1,196 @@
+import asyncio
+import dataclasses
+from copy import deepcopy
+from typing import Any, Dict, List, Literal
+
+from agenthub.gptswarm_agent.gptswarm_graph import AssistantGraph
+from agenthub.gptswarm_agent.prompt import GPTSwarmPromptSet
+from opendevin.controller.agent import Agent
+from opendevin.controller.state.state import State
+from opendevin.core.logger import opendevin_logger as logger
+from opendevin.events.action import Action
+from opendevin.llm.llm import LLM
+
+ENABLE_GITHUB = True
+OPENAI_API_KEY = 'sk-proj-****'  # TODO: get from environment or config
+
+
+MessageRole = Literal['system', 'user', 'assistant']
+
+
+@dataclasses.dataclass()
+class Message:
+    role: MessageRole
+    content: str
+
+
+class GPTSwarm(Agent):
+    VERSION = '1.0'
+    """
+    This is simple revision of GPTSwarm which serve as an assistant agent.
+
+    GPTSwarm Paper: https://arxiv.org/abs/2402.16823 (ICML 2024, Oral Presentation)
+    GPTSwarm Code: https://github.com/metauto-ai/GPTSwarm
+    """
+
+    def __init__(
+        self,
+        llm: LLM,
+        model_name: str,
+    ) -> None:
+        """
+        Initializes a new instance of the GPTSwarm class.
+
+        Parameters:
+        - llm (LLM): The llm to be used by this agent
+        """
+        super().__init__(llm)
+        self.api_key = OPENAI_API_KEY
+        self.llm = LLM(model=model_name, api_key=self.api_key)
+        self.graph = AssistantGraph(domain='gaia', model_name=model_name)
+        self.prompt_set = GPTSwarmPromptSet()
+
+    def reset(self) -> None:
+        """
+        Resets the GPTSwarm Agent.
+        """
+        super().reset()
+
+    def step(self, state: State) -> Action:
+        """
+        # TODO: It is stateless now. Find a way to make it stateful.
+        # NOTE: For the AI assistant, state-based design may introduce more uncertainties.
+        """
+        raise NotImplementedError
+
+    async def swarm_run(self, inputs: List[Dict[str, Any]], num_agents=3) -> List[str]:
+        """
+        Run the `run` method of this agent concurrently for `num_agents` times.
+        # NOTE: This is just a simple self-consistency.
+        # TODO: should follow original GPTSwarm's graph design to revise.
+        """
+
+        async def run_single_agent(index):
+            try:
+                result = await asyncio.wait_for(self.run(inputs=inputs), timeout=200)
+                print('-----------------------------------')
+                print(f'No. {index} Agent complete task..')
+                logger.info(result[0])
+                print('-----------------------------------')
+                return result[0]
+            except asyncio.TimeoutError:
+                print(f'No. {index} Agent timed out.')
+                return None
+            except Exception as e:
+                print(f'No. {index} Agent resulted in an error: {e}')
+                return None
+
+        # Create a list of tasks to run concurrently
+        tasks = [run_single_agent(i) for i in range(num_agents)]
+
+        # Run all tasks concurrently and gather the results
+        agent_answers = await asyncio.gather(*tasks)
+
+        # Filter out None results (from timeouts or errors)
+        agent_answers = [answer for answer in agent_answers if answer is not None]
+
+        task = inputs[0]['task']
+        prompt = self.prompt_set.get_self_consistency(
+            question=task,
+            answers=agent_answers,
+            constraint=self.prompt_set.get_constraint(),
+        )
+        messages = [
+            Message(role='system', content=f'You are a {self.prompt_set.get_role()}.'),
+            Message(role='user', content=prompt),
+        ]
+
+        swarm_ans = self.llm.completion(
+            messages=[{'role': msg.role, 'content': msg.content} for msg in messages]
+        )
+        swarm_ans = swarm_ans.choices[0].message.content
+        return [swarm_ans]
+
+    async def run(
+        self,
+        inputs: List[Dict[str, Any]],
+        max_tries: int = 3,
+        max_time: int = 600,
+        return_all_outputs: bool = False,
+    ) -> List[Any]:
+        def is_node_useful(node):
+            if node in self.graph.output_nodes:
+                return True
+
+            for successor in node.successors:
+                if is_node_useful(successor):
+                    return True
+            return False
+
+        useful_node_ids = [
+            node_id
+            for node_id, node in self.graph.nodes.items()
+            if is_node_useful(node)
+        ]
+        in_degree = {
+            node_id: len(self.graph.nodes[node_id].predecessors)
+            for node_id in useful_node_ids
+        }
+        zero_in_degree_queue = [
+            node_id
+            for node_id, deg in in_degree.items()
+            if deg == 0 and node_id in useful_node_ids
+        ]
+
+        for i, input_node in enumerate(self.graph.input_nodes):
+            node_input = deepcopy(inputs)
+            input_node.inputs = [node_input]
+
+        while zero_in_degree_queue:
+            current_node_id = zero_in_degree_queue.pop(0)
+            current_node = self.graph.nodes[current_node_id]
+            tries = 0
+            while tries < max_tries:
+                try:
+                    await asyncio.wait_for(
+                        self.graph.nodes[current_node_id].execute(), timeout=max_time
+                    )
+                    # TODO: make GPTSwarm stateful in OpenDevin.
+                    # State.inputs = self.graph.nodes[current_node_id].inputs
+                    # State.outputs = self.graph.nodes[current_node_id].outputs
+                    # self.step(State)
+
+                except asyncio.TimeoutError:
+                    print(
+                        f'Node {current_node_id} execution timed out, retrying {tries + 1} out of {max_tries}...'
+                    )
+                except Exception as e:
+                    print(f'Error during execution of node {current_node_id}: {e}')
+                    break
+                tries += 1
+
+            for successor in current_node.successors:
+                if successor.id in useful_node_ids:
+                    in_degree[successor.id] -= 1
+                    if in_degree[successor.id] == 0:
+                        zero_in_degree_queue.append(successor.id)
+
+        final_answers = []
+
+        for output_node in self.graph.output_nodes:
+            output_messages = output_node.outputs
+
+            if len(output_messages) > 0 and not return_all_outputs:
+                final_answer = output_messages[-1].get('output', output_messages[-1])
+                final_answers.append(final_answer)
+            else:
+                for output_message in output_messages:
+                    final_answer = output_message.get('output', output_message)
+                    final_answers.append(final_answer)
+
+        if len(final_answers) == 0:
+            final_answers.append('No answer since there are no inputs provided')
+        return final_answers
+
+    def search_memory(self, query: str) -> list[str]:
+        raise NotImplementedError('Implement this abstract method')
--- a/agenthub/gptswarm_agent/gptswarm_graph.py
+++ b/agenthub/gptswarm_agent/gptswarm_graph.py
@@ -0,0 +1,520 @@
+#!/usr/bin/env python
+# -*- coding: utf-8 -*-
+
+import ast
+import asyncio
+import dataclasses
+import os
+import re
+from collections import defaultdict
+from pathlib import Path
+from typing import Any, List, Literal, Optional
+
+import requests
+from pytube import YouTube
+from swarm.graph import Graph, Node
+
+from agenthub.gptswarm_agent.prompt import GPTSwarmPromptSet
+from opendevin.core.logger import opendevin_logger as logger
+from opendevin.llm.llm import LLM
+from opendevin.runtime.plugins.agent_skills.agentskills import (
+    parse_audio,
+    parse_docx,
+    parse_image,
+    parse_latex,
+    parse_pdf,
+    parse_pptx,
+    parse_txt,
+    parse_video,
+)
+
+OPENAI_API_KEY = 'sk-proj-****'  # TODO: get from environment or config
+SEARCHAPI_API_KEY = '****'  # TODO: get from environment or config
+
+MessageRole = Literal['system', 'user', 'assistant']
+
+
+@dataclasses.dataclass()
+class Message:
+    role: MessageRole
+    content: str
+
+
+READER_MAP = {
+    '.png': parse_image,
+    '.jpg': parse_image,
+    '.jpeg': parse_image,
+    '.gif': parse_image,
+    '.bmp': parse_image,
+    '.tiff': parse_image,
+    '.tif': parse_image,
+    '.webp': parse_image,
+    '.mp3': parse_audio,
+    '.m4a': parse_audio,
+    '.wav': parse_audio,
+    '.MOV': parse_video,
+    '.mp4': parse_video,
+    '.mov': parse_video,
+    '.avi': parse_video,
+    '.mpg': parse_video,
+    '.mpeg': parse_video,
+    '.wmv': parse_video,
+    '.flv': parse_video,
+    '.webm': parse_video,
+    '.pptx': parse_pptx,
+    '.pdf': parse_pdf,
+    '.docx': parse_docx,
+    '.tex': parse_latex,
+    '.txt': parse_txt,
+}
+
+
+class FileReader:
+    def __init__(self):
+        self.reader = None  # Initial type is None
+
+    def set_reader(self, suffix: str):
+        reader = READER_MAP.get(suffix)
+        if reader is not None:
+            self.reader = reader
+            logger.info(f'Setting Reader to {self.reader.__name__}')
+        else:
+            logger.error(f'No reader found for suffix {suffix}')
+            self.reader = None
+
+    def read_file(self, file_path: Path, task: str = 'describe the file') -> str:
+        suffix = file_path.suffix
+        self.set_reader(suffix)
+        if not self.reader:
+            raise ValueError(f'No reader set for suffix {suffix}')
+        if self.reader in [parse_image, parse_video]:
+            file_content = self.reader(file_path, task)
+        else:
+            file_content = self.reader(file_path)
+        logger.info(f'Reading file {file_path} using {self.reader.__name__}')
+        return file_content
+
+
+class GenerateQuery(Node):
+    def __init__(
+        self,
+        domain: str = 'gaia',
+        model_name: Optional[str] = 'gpt-4o-2024-05-13',
+        operation_description: str = 'Given a question, return what information is needed to answer the question.',
+        id=None,
+    ):
+        super().__init__(operation_description, id, True)
+        self.domain = domain
+        self.api_key = OPENAI_API_KEY
+        self.llm = LLM(model=model_name, api_key=self.api_key)
+        self.prompt_set = GPTSwarmPromptSet()
+
+    @property
+    def node_name(self) -> str:
+        return self.__class__.__name__
+
+    def extract_urls(self, text: str) -> List[str]:
+        url_pattern = r'https?://[^\s]+'
+        urls = re.findall(url_pattern, text)
+        return urls
+
+    def is_youtube_url(self, url: str) -> bool:
+        youtube_regex = (
+            r'(https?://)?(www\.)?'
+            r'(youtube|youtu|youtube-nocookie)\.(com|be)/'
+            r'(watch\?v=|embed/|v/|.+\?v=)?([^&=%\?]{11})'
+        )
+        return bool(re.match(youtube_regex, url))
+
+    def _youtube_download(self, url: str) -> str:
+        try:
+            video_id = url.split('v=')[-1].split('&')[0]
+            video_id = video_id.strip()
+            youtube = YouTube(url)
+            video_stream = (
+                youtube.streams.filter(progressive=True, file_extension='mp4')
+                .order_by('resolution')
+                .desc()
+                .first()
+            )
+            if not video_stream:
+                raise ValueError('No suitable video stream found.')
+
+            output_dir = 'workspace/tmp'
+            os.makedirs(output_dir, exist_ok=True)
+            output_path = f'{output_dir}/{video_id}.mp4'
+            video_stream.download(output_path=output_dir, filename=f'{video_id}.mp4')
+            return output_path
+
+        except Exception as e:
+            logger.error(
+                f'Error downloading video from {url}: {e}'
+            )  # Use logger for error messages
+            return ''
+
+    async def _execute(
+        self, inputs: Optional[List[dict]] = None, **kwargs
+    ) -> List[dict]:
+        if inputs is None:
+            inputs = []
+        node_inputs = inputs
+        outputs = []
+
+        for input in node_inputs:
+            urls = self.extract_urls(input['task'])
+
+            download_paths = []
+
+            for url in urls:
+                if self.is_youtube_url(url):
+                    download_path = self._youtube_download(url)
+                    if download_path:
+                        download_paths.append(download_path)
+
+            if urls:
+                logger.info(urls)
+            if download_paths:
+                logger.info(download_paths)
+
+            files = input.get('files', [])
+            if not isinstance(files, list):
+                files = []
+            files.extend(download_paths)
+
+            role = self.prompt_set.get_role()
+            # constraint = self.prompt_set.get_constraint()
+            prompt = self.prompt_set.get_query_prompt(question=input['task'])
+
+            messages = [
+                Message(role='system', content=f'You are a {role}.'),
+                Message(role='user', content=prompt),
+            ]
+
+            response = self.llm.completion(
+                messages=[
+                    {'role': msg.role, 'content': msg.content} for msg in messages
+                ]
+            )
+            response = response.choices[0].message.content
+
+            executions = {
+                'operation': self.node_name,
+                'task': input['task'],
+                'files': files,
+                'input': input.get('task', None),
+                'subtask': prompt,
+                'output': response,
+                'format': 'natural language',
+            }
+            outputs.append(executions)
+
+        return outputs
+
+
+class FileAnalyse(Node):
+    def __init__(
+        self,
+        domain: str = 'gaia',
+        model_name: Optional[str] = 'gpt-4o-2024-05-13',
+        operation_description: str = 'Given a question, extract information from a file.',
+        id=None,
+    ):
+        super().__init__(operation_description, id, True)
+        self.domain = domain
+        self.api_key = OPENAI_API_KEY
+        self.llm = LLM(model=model_name, api_key=self.api_key)
+        self.prompt_set = GPTSwarmPromptSet()
+        self.reader = FileReader()
+
+    @property
+    def node_name(self) -> str:
+        return self.__class__.__name__
+
+    async def _execute(
+        self, inputs: Optional[List[dict]] = None, **kwargs
+    ) -> List[dict]:
+        if inputs is None:
+            inputs = []
+        node_inputs = inputs
+        outputs = []
+        for input in node_inputs:
+            query = input.get('output', 'Please organize the information of this file.')
+            files = input.get('files', [])
+            response = await self.file_analyse(query, files, self.llm)
+
+            executions = {
+                'operation': self.node_name,
+                'task': input['task'],
+                'files': files,
+                'input': query,
+                'subtask': f'Read the content of ###{files}, use query ###{query}',
+                'output': response,
+                'format': 'natural language',
+            }
+
+            outputs.append(executions)
+
+        return outputs
+
+    async def file_analyse(self, query: str, files: List[str], llm: LLM) -> str:
+        answer = ''
+        for file in files:
+            file_path = Path(file)
+            if self.reader not in [parse_image, parse_video]:
+                file_content = self.reader.read_file(file_path)
+                prompt = self.prompt_set.get_file_analysis_prompt(
+                    query=query, file=file_content
+                )
+                messages = [
+                    Message(
+                        role='system',
+                        content=f'You are a {self.prompt_set.get_role()}.',
+                    ),
+                    Message(role='user', content=prompt),
+                ]
+                response = llm.completion(
+                    messages=[
+                        {'role': msg.role, 'content': msg.content} for msg in messages
+                    ]
+                )
+                answer += response.choices[0].message.content + '\n'
+        return answer
+
+
+class WebSearch(Node):
+    def __init__(
+        self,
+        domain: str = 'gaia',
+        model_name: Optional[str] = 'gpt-4o-2024-05-13',
+        operation_description: str = 'Given a question, search the web for infomation.',
+        id=None,
+    ):
+        super().__init__(operation_description, id, True)
+        self.domain = domain
+        self.api_key = OPENAI_API_KEY
+        self.llm = LLM(model=model_name, api_key=self.api_key)
+        self.prompt_set = GPTSwarmPromptSet()
+
+    @property
+    def node_name(self) -> str:
+        return self.__class__.__name__
+
+    async def _execute(
+        self, inputs: Optional[List[dict]] = None, max_keywords: int = 4, **kwargs
+    ) -> List[dict]:
+        if inputs is None:
+            inputs = []
+        node_inputs = inputs
+        outputs = []
+        for input in node_inputs:
+            task = input['task']
+            query = input['output']
+            prompt = self.prompt_set.get_websearch_prompt(question=task, query=query)
+            messages = [
+                Message(
+                    role='system', content=f'You are a {self.prompt_set.get_role()}.'
+                ),
+                Message(role='user', content=prompt),
+            ]
+            generated_quires = self.llm.completion(
+                messages=[
+                    {'role': msg.role, 'content': msg.content} for msg in messages
+                ]
+            )
+
+            generated_quires = generated_quires.choices[0].message.content
+            generated_quires = generated_quires.split(',')[:max_keywords]
+            logger.info(f'The search keywords include: {generated_quires}')
+            search_results = [self.web_search(query) for query in generated_quires]
+            logger.info(f'The search results: {str(search_results)[:100]}...')
+
+            distill_prompt = self.prompt_set.get_distill_websearch_prompt(
+                question=input['task'], query=query, results='.\n'.join(search_results)
+            )
+
+            messages = [
+                Message(
+                    role='system', content=f'You are a {self.prompt_set.get_role()}.'
+                ),
+                Message(role='user', content=distill_prompt),
+            ]
+            response = self.llm.completion(
+                messages=[
+                    {'role': msg.role, 'content': msg.content} for msg in messages
+                ]
+            )
+            response = response.choices[0].message.content
+
+            executions = {
+                'operation': self.node_name,
+                'task': task,
+                'files': input.get('files', []),
+                'input': query,
+                'subtask': distill_prompt,
+                'output': response,
+                'format': 'natural language',
+            }
+            outputs.append(executions)
+
+        return outputs
+
+    def web_search(self, query: str, item_num: int = 3) -> str:
+        url = 'https://www.searchapi.io/api/v1/search'
+        params = {
+            'engine': 'google',
+            'q': query,
+            'api_key': SEARCHAPI_API_KEY,  # os.getenv("SEARCHAPI_API_KEY")
+        }
+
+        response = ast.literal_eval(requests.get(url, params=params).text)
+
+        if (
+            'knowledge_graph' in response.keys()
+            and 'description' in response['knowledge_graph'].keys()
+        ):
+            return response['knowledge_graph']['description']
+
+        if (
+            'organic_results' in response.keys()
+            and len(response['organic_results']) > 0
+        ):
+            snippets = []
+            for res in response['organic_results'][:item_num]:
+                if 'snippet' in res:
+                    snippets.append(res['snippet'])
+            return '\n'.join(snippets)
+
+        return ' '
+
+
+class CombineAnswer(Node):
+    def __init__(
+        self,
+        domain: str = 'gaia',
+        model_name: Optional[str] = 'gpt-4o-2024-05-13',
+        operation_description: str = 'Combine multiple inputs into one.',
+        max_token: int = 500,
+        id=None,
+    ):
+        super().__init__(operation_description, id, True)
+        self.domain = domain
+        self.max_token = max_token
+        self.api_key = OPENAI_API_KEY
+        self.llm = LLM(model=model_name, api_key=self.api_key)
+        self.prompt_set = GPTSwarmPromptSet()
+        self.materials: defaultdict[str, str] = defaultdict(str)
+
+    @property
+    def node_name(self) -> str:
+        return self.__class__.__name__
+
+    async def _execute(
+        self, inputs: Optional[List[Any]] = None, **kwargs
+    ) -> List[dict]:
+        if inputs is None:
+            inputs = []
+        node_inputs = inputs
+
+        role = self.prompt_set.get_role()
+        constraint = self.prompt_set.get_constraint()
+
+        self.materials = defaultdict(str)
+        for input in node_inputs:
+            operation = input.get('operation')
+            if operation:
+                self.materials[operation] += f'{input.get("output", "")}\n'
+            self.materials['task'] = input.get('task')
+
+        question = self.prompt_set.get_combine_materials(self.materials)
+        prompt = self.prompt_set.get_answer_prompt(question=question)
+
+        messages = [
+            Message(role='system', content=f'You are a {role}. {constraint}'),
+            Message(role='user', content=prompt),
+        ]
+
+        response = self.llm.completion(
+            messages=[{'role': msg.role, 'content': msg.content} for msg in messages]
+        )
+
+        response = response.choices[0].message.content
+
+        executions = {
+            'operation': self.node_name,
+            'task': self.materials['task'],
+            'files': self.materials['files']
+            if isinstance(self.materials['files'], str)
+            else ', '.join(self.materials['files']),
+            'input': node_inputs,
+            'subtask': prompt,
+            'output': response,
+            'format': 'natural language',
+        }
+
+        return [executions]
+
+
+class AssistantGraph(Graph):
+    def build_graph(self):
+        query = GenerateQuery(self.domain, self.model_name)
+
+        file_analysis = FileAnalyse(self.domain, self.model_name)
+        web_search = WebSearch(self.domain, self.model_name)
+
+        query.add_successor(file_analysis)
+        query.add_successor(web_search)
+
+        combine = CombineAnswer(self.domain, self.model_name)
+        file_analysis.add_successor(combine)
+        web_search.add_successor(combine)
+
+        self.input_nodes = [query]
+        self.output_nodes = [combine]
+
+        self.add_node(query)
+        self.add_node(file_analysis)
+        self.add_node(web_search)
+        self.add_node(combine)
+
+
+if __name__ == '__main__':
+    # # test node
+    # task = 'What is the text representation of the last digit of twelve squared?'
+    # inputs = [{'task': task}]
+    # query_instance = GenerateQuery()
+    # query = asyncio.run(query_instance._execute(inputs))
+    # print(query)
+
+    # task = 'What is the text representation of the last digit of twelve squared?'
+    # inputs = [
+    #     {
+    #         'task': 'How can researchers ensure AGI development is both safe and ethical while avoiding societal biases and inequalities?',
+    #         'files': ['agi.txt'],
+    #     }
+    # ]
+    # file_instance = FileAnalyse()
+    # file_info = asyncio.run(file_instance._execute(inputs))
+    # print(file_info)
+
+    # task = 'What is the text representation of the last digit of twelve squared?'
+    # inputs = [
+    #     {
+    #         'task': 'How can researchers ensure AGI development is both safe and ethical while avoiding societal biases and inequalities?'
+    #     }
+    # ]
+    # search_instance = WebSearch()
+    # search_info = asyncio.run(search_instance._execute(inputs))
+    # print(search_info)
+
+    assistant_graph = AssistantGraph(domain='gaia', model_name='gpt-4o-2024-05-13')
+
+    # test graph
+    assistant_graph.build_graph()
+    inputs = [
+        {
+            'task': 'How can researchers ensure AGI development is both safe and ethical while avoiding societal biases and inequalities?',
+            'files': ['agi.txt'],
+        }
+    ]
+    outputs = asyncio.run(assistant_graph.run(inputs))
+    print(outputs)
--- a/agenthub/gptswarm_agent/prompt.py
+++ b/agenthub/gptswarm_agent/prompt.py
@@ -0,0 +1,129 @@
+#!/usr/bin/env python
+# -*- coding: utf-8 -*-
+
+from typing import Any, Dict
+
+
+class GPTSwarmPromptSet:
+    """
+    GPTSwarmPromptSet provides a collection of static methods to generate prompts
+    for a general AI assistant. These prompts cover various tasks like answering questions,
+    performing web searches, analyzing files, and reflecting on tasks.
+    """
+
+    @staticmethod
+    def get_role():
+        return 'a general AI assistant'
+
+    @staticmethod
+    def get_constraint():
+        return (
+            'I will ask you a question. Report your thoughts, and finish your answer with the following template: FINAL ANSWER: [YOUR FINAL ANSWER]. '
+            'YOUR FINAL ANSWER should be a number OR as few words as possible OR a comma separated list of numbers and/or strings. '
+            "If you are asked for a number, don't use comma to write your number neither use units such as $ or percent sign unless specified otherwise. "
+            "If you are asked for a string, don't use articles, neither abbreviations (e.g. for cities), and write the digits in plain text unless specified otherwise. "
+            'If you are asked for a comma separated list, apply the above rules depending of whether the element to be put in the list is a number or a string. '
+        )
+
+    @staticmethod
+    def get_format():
+        return 'natural language'
+
+    @staticmethod
+    def get_answer_prompt(question):
+        return f'{question}'
+
+    @staticmethod
+    def get_query_prompt(question):
+        return (
+            '# Information Gathering for Question Resolution\n\n'
+            'Evaluate if additional information is needed to answer the question. '
+            'If a web search or file analysis is necessary, outline specific clues or details to be searched for.\n\n'
+            f'## ❓ Target Question:\n{question}\n\n'
+            '## 🔍 Clues for Investigation:\n'
+            'Identify critical clues and concepts within the question that are essential for finding the answer.\n'
+        )
+
+    @staticmethod
+    def get_file_analysis_prompt(query, file):
+        return (
+            '# File Analysis Task\n\n'
+            f'## 🔍 Information Extraction Objective:\n---\n{query}\n---\n\n'
+            f'## 📄 File Under Analysis:\n---\n{file}\n---\n\n'
+            '## 📝 Instructions:\n'
+            '1. Identify the key sections in the file relevant to the query.\n'
+            '2. Extract and summarize the necessary information from these sections.\n'
+            '3. Ensure the response is focused and directly addresses the query.\n'
+            "Example: 'Identify the main theme in the text.'"
+        )
+
+    @staticmethod
+    def get_websearch_prompt(question, query):
+        return (
+            '# Web Search Task\n\n'
+            f'## Original Question: \n---\n{question}\n---\n\n'
+            f'## 🔍 Targeted Search Objective:\n---\n{query}\n---\n\n'
+            '## 🌐 Simplified Search Instructions:\n'
+            'Generate three specific search queries directly related to the original question. Each query should focus on key terms from the question. Format the output as a comma-separated list.\n'
+            "For example, if the question is 'Who will be the next US president?', your queries could be: 'US presidential candidates, current US president, next US president'.\n"
+            "Remember to format the queries as 'query1, query2, query3'."
+        )
+
+    @staticmethod
+    def get_distill_websearch_prompt(question, query, results):
+        return (
+            '# Summarization of Search Results\n\n'
+            f'## Original question: \n---\n{question}\n---\n\n'
+            f'## 🔍 Required Information for Summary:\n---\n{query}\n---\n\n'
+            f'## 🌐 Analyzed Search Results:\n---\n{results}\n---\n\n'
+            '## 📝 Instructions for Summarization:\n'
+            '1. Review the provided search results and identify the most relevant information related to the question and query.\n'
+            '2. Extract and highlight the key findings, facts, or data points from these results.\n'
+            '3. Organize the summarized information in a coherent and logical manner.\n'
+            '4. Ensure the summary is concise and directly addresses the query, avoiding extraneous details.\n'
+            '5. If the information from web search is useless, directly answer: "No useful information from WebSearch".\n'
+        )
+
+    @staticmethod
+    def get_combine_materials(materials: Dict[str, Any], avoid_vague=True) -> str:
+        question = materials.get('task', 'No problem provided')
+
+        for key, value in materials.items():
+            if 'No useful information from WebSearch' in value:
+                continue
+            value = value.strip('\n').strip()
+            if key != 'task' and value:
+                question += (
+                    f'\n\nReference information for {key}:'
+                    + '\n----------------------------------------------\n'
+                    + f'{value}'
+                    + '\n----------------------------------------------\n\n'
+                )
+
+        if avoid_vague:
+            question += (
+                '\nProvide a specific answer. For questions with known answers, ensure to provide accurate and factual responses. '
+                + "Avoid vague responses or statements like 'unable to...' that don't contribute to a definitive answer. "
+                + "For example: if a question asks 'who will be the president of America', and the answer is currently unknown, you could suggest possibilities like 'Donald Trump', or 'Biden'. However, if the answer is known, provide the correct information."
+            )
+
+        return question
+
+    @staticmethod
+    def get_self_consistency(question: str, answers: list, constraint: str) -> str:
+        formatted_answers = '\n'.join(
+            [f'Answer {index + 1}: {answer}' for index, answer in enumerate(answers)]
+        )
+        return (
+            '# Self-Consistency Evaluation Task\n\n'
+            f'## 🤔 Question for Review:\n---\n{question}\n---\n\n'
+            f'## 💡 Reviewable Answers:\n---\n{formatted_answers}\n---\n\n'
+            '## 📋 Instructions for Selection:\n'
+            '1. Read each answer and assess how it addresses the question.\n'
+            "2. Compare the answers for their adherence to the given question's criteria and logical coherence.\n"
+            "3. Identify the answer that best aligns with the question's requirements and is the most logically consistent.\n"
+            "4. Ignore the candidate answers if they do not give a direct answer, for example, using 'unable to ...', 'as an AI ...'.\n"
+            '5. Copy the most suitable answer as it is, without modification, to maintain its original form.\n'
+            f'6. Adhere to the constraints: {constraint}.\n'
+            'Note: If no answer fully meets the criteria, choose and copy the one that is closest to the requirements.'
+        )
--- a/agenthub/micro/README.md
+++ b/agenthub/micro/README.md
@@ -0,0 +1,17 @@
+## Introduction
+
+This package contains definitions of micro-agents. A micro-agent is defined
+in the following structure:
+
+```
+[AgentName]
+├── agent.yaml
+└── prompt.md
+```
+
+Note that `prompt.md` could use jinja2 template syntax. During runtime, `prompt.md`
+is loaded and rendered, and used together with `agent.yaml` to initialize a
+micro-agent.
+
+Micro-agents can be used independently. You can also use `ManagerAgent` which knows
+how to coordinate the agents and collaboratively finish a task.
--- a/agenthub/micro/_instructions/actions/browse.md
+++ b/agenthub/micro/_instructions/actions/browse.md
@@ -0,0 +1,2 @@
+* `browse` - opens a web page. Arguments:
+  * `url` - the URL to open
--- a/Show More
+++ b/Show More
				`@@ -1 +0,0 @@`
				`This way of running OpenHands is not officially supported. It is maintained by the community.`