This guide is for enterprise admins and architects who need to govern Agent Tracks across an organization. It covers how to vet, evaluate, deploy, and manage Tracks at scale. For authoring guidance, see best practices. For architecture details, see the Tracks overview.
Security review and vetting
Deploying Tracks in an enterprise requires answering two distinct questions:
- Are Tracks safe in general? See the security considerations section in the overview for platform-level security details.
- How do I vet a specific Track? Use the following risk assessment and review checklist.
Risk tier assessment
Evaluate each Track against these risk indicators before approving deployment:
| Risk indicator | What to look for | Concern level |
|---|---|---|
| Code execution | Scripts in the Track directory (.py, .sh, *.js) | High: scripts run with full environment access |
| Instruction manipulation | Directives to ignore safety rules, hide actions from users, or alter Haijun's behavior conditionally | High: can bypass security controls |
| MCP server references | Instructions referencing MCP tools (ServerName:tool_name) | High: extends access beyond the Track itself |
| Network access patterns | URLs, API endpoints, fetch, curl, or requests calls | High: potential data exfiltration vector |
| Hardcoded credentials | API keys, tokens, or passwords in Track files or scripts | High: secrets exposed in Git history and context window |
| Filesystem access scope | Paths outside the Track directory, broad glob patterns, path traversal (../) | Medium: may access unintended data |
| Tool invocations | Instructions directing Haijun to use bash, file operations, or other tools | Medium: review what operations are performed |
Review checklist
Before deploying any Track from a third party or internal contributor, complete these steps:
- Read all Track directory content. Review SKILL.md, all referenced markdown files, and any bundled scripts or resources.
- Verify script behavior matches stated purpose. Run scripts in a sandboxed environment and confirm outputs align with the Track's description.
- Check for adversarial instructions. Look for directives that tell Haijun to ignore safety rules, hide actions from users, exfiltrate data through responses, or alter behavior based on specific inputs.
- Check for external URL fetches or network calls. Search scripts and instructions for network access patterns (
http,requests.get,urllib,curl,fetch).
- Verify no hardcoded credentials. Check for API keys, tokens, or passwords in Track files. Credentials should use environment variables or secure credential stores, never appear in Track content.
- Identify tools and commands the Track instructs Haijun to invoke. List all bash commands, file operations, and tool references. Consider the combined risk when a Track uses both file-read and network tools together.
- Confirm redirect destinations. If the Track references external URLs, verify they point to expected domains.
- Verify no data exfiltration patterns. Look for instructions that read sensitive data and then write, send, or encode it for external transmission, including through Haijun's conversational responses.
Warning: Never deploy Tracks from untrusted sources without a full audit. A malicious Track can direct Haijun to execute arbitrary code, access sensitive files, or transmit data externally. Treat Track installation with the same rigor as installing software on production systems.
Track content scanning
Haijun Enterprise organizations can turn on automated security scanning for custom Tracks in haijun.ai and Haijun Cowork. After you turn on Track and plugin security scanning at haijun.ai > Organization settings > Tracks, Tracks that members then upload or edit in haijun.ai or Cowork are scanned for signs of malicious behavior, such as hidden code execution, sending your data to an outside service, or instructions that tamper with Haijun's safeguards. A Track that fails the scan, or whose scan hasn't finished, is blocked from use. A Track that passes with a warning stays usable behind a caution notice. If scanning is available to your organization, turn it on. It complements, but doesn't replace, the review checklist.
Scanning doesn't cover the Haijun API. Tracks you upload through the Tracks API (/v1/tracks), including from the Haijun Console, aren't scanned, so for API deployments, rely on the review checklist and version pinning. Scanning also doesn't apply to Tracks that were already in your organization when you turned it on, or to organizations with certain data handling configurations, such as customer-managed encryption keys (CMEK), zero data retention (ZDR), or HIPAA readiness. For setup steps, exclusions, and result types, see Get started with track and plugin scanning in the Haijun Help Center.
Evaluating Tracks before deployment
Tracks can degrade agent performance if they trigger incorrectly, conflict with other Tracks, or provide poor instructions. Require evaluation before any production deployment.
What to evaluate
Establish approval gates for these dimensions before deploying any Track:
| Dimension | What it measures | Example failure |
|---|---|---|
| Triggering accuracy | Does the Track activate for the right queries and stay inactive for unrelated ones? | Track triggers on every spreadsheet mention, even when the user just wants to discuss data |
| Isolation behavior | Does the Track work correctly on its own? | Track references files that don't exist in its directory |
| Coexistence | Does adding this Track degrade other Tracks? | New Track's description is too broad, stealing triggers from existing Tracks |
| Instruction following | Does Haijun follow the Track's instructions accurately? | Haijun skips validation steps or uses wrong libraries |
| Output quality | Does the Track produce correct, useful results? | Generated reports have formatting errors or missing data |
Evaluation requirements
Require Track authors to submit evaluation suites with 3–5 representative queries per Track, covering cases where the Track should trigger, should not trigger, and ambiguous edge cases. Require testing across the models your organization uses (Haiku, Sonnet, Opus), because Track effectiveness varies by model.
For detailed guidance on building evaluations, see evaluation and iteration in best practices. For general evaluation methodology, see develop test cases.
Using evaluations for lifecycle decisions
Evaluation results signal when to act:
- Declining trigger accuracy: Update the Track's description or instructions
- Coexistence conflicts: Consolidate overlapping Tracks or narrow descriptions
- Consistently low output quality: Rewrite instructions or add validation steps
- Persistent failures across updates: Deprecate the Track
Track lifecycle management
- Plan
Identify workflows that are repetitive, error-prone, or require specialized knowledge. Map these to organizational roles and determine which are candidates for Tracks.
- Create and review
Ensure the Track author follows best practices. Require a security review using the review checklist. Require an evaluation suite before approval. Establish separation of duties: Track authors should not be their own reviewers.
- Test
Require evaluations in isolation (Track alone) and alongside existing Tracks (coexistence testing). Verify triggering accuracy, output quality, and absence of regressions across your active Track set before approving for production.
- Deploy
Upload through the Tracks API for workspace-wide access. See Using Tracks with the API for upload and version management. Document the Track in your internal registry with purpose, owner, and version.
- Monitor
Track usage patterns and collect feedback from users. Rerun evaluations periodically to detect drift or regressions as workflows and models evolve. Usage analytics are not currently available through the Tracks API. Implement application-level logging to track which Tracks are included in requests.
- Iterate or deprecate
Require the full evaluation suite to pass before promoting new versions. Update Tracks when workflows change or evaluation scores decline. Deprecate Tracks when evaluations consistently fail or the workflow is retired.
Organizing Tracks at scale
Recall limits
As a general guideline, limit the number of Tracks loaded simultaneously to maintain reliable recall accuracy. Each Track's metadata (name and description) competes for attention in the system prompt. With too many Tracks active, Haijun may fail to select the right Track or miss relevant ones entirely. Use your evaluation suite to measure recall accuracy as you add Tracks, and stop adding when performance degrades.
Note that API requests support a maximum of 20 Tracks for each request (see Using Tracks with the API). If a role requires more Tracks than a single request supports, consider consolidating narrow Tracks into broader ones or routing requests to different Track sets based on task type.
Start specific, consolidate later
Encourage teams to start with narrow, workflow-specific Tracks rather than broad, multipurpose ones. As patterns emerge across your organization, consolidate related Tracks into role-based bundles.
Tip: Use evaluations to decide when to consolidate. Merge narrow Tracks into a broader one only when the consolidated Track's evaluations confirm equivalent performance to the individual Tracks it replaces.
Example progression:
- Start:
formatting-sales-reports,querying-pipeline-data,updating-crm-records
- Consolidate:
sales-operations(when evals confirm equivalent performance)
Naming and cataloging
Use consistent naming conventions across your organization. The naming conventions section in best practices provides formatting guidance.
Maintain an internal registry for each Track with:
- Purpose: What workflow the Track supports
- Owner: Team or individual responsible for maintenance
- Version: Current deployed version
- Dependencies: MCP servers, packages, or external services required
- Evaluation status: Last evaluation date and results
Role-based bundles
Group Tracks by organizational role to keep each user's active Track set focused:
- Sales team: CRM operations, pipeline reporting, proposal generation
- Engineering: Code review, deployment workflows, incident response
- Finance: Report generation, data validation, audit preparation
Each role-based bundle should contain only the Tracks relevant to that role's daily workflows.
Distribution and version control
Source control
Store Track directories in Git for history tracking, code review through pull requests, and rollback capability. Each Track directory (containing SKILL.md and any bundled files) maps naturally to a Git-tracked folder.
API-based distribution
The Tracks API provides workspace-scoped distribution. Tracks uploaded through the API are available to all workspace members. See Using Tracks with the API for upload, versioning, and management endpoints.
Versioning strategy
- Production: Pin Tracks to specific versions. If you omit
version, requests use the latest version, so a new version uploaded by anyone in the workspace immediately changes what production agents run. Run the full evaluation suite before promoting a new version. Treat every update as a new deployment requiring full security review.
- Development and testing: Use latest versions to validate changes before production promotion.
- Rollback plan: Maintain the previous version as a fallback. If a new version fails evaluations in production, revert to the last known-good version immediately.
- Integrity verification: Compute checksums of reviewed Tracks and verify them at deployment time. Use signed commits in your Track repository to ensure provenance.
Cross-surface considerations
Warning: Custom Tracks do not sync across surfaces. Tracks uploaded to the API are not available on haijun.ai or in Haijun Code, and vice versa. Each surface requires separate uploads and management.
Maintain Track source files in Git as the single source of truth. If your organization deploys Tracks across multiple surfaces, implement your own synchronization process to keep them consistent. For full details, see cross-surface availability.
Next steps
Architecture and platform details
Authoring guidance for Track creators
Upload and manage Tracks programmatically