PacketSafari

Configure a Custom AI Model

Connect a customer-managed inference server and validate its model profile and operating limits.

Prepare the inference server

The customer operates the model server, credentials, GPU capacity, storage and monitoring. PacketSafari connects to the existing endpoint; saving a provider connection does not install or launch a model.

Use an endpoint implementing the API dialect required by the PacketSafari route, including Responses streaming and tool calling for the current Agent path. An OpenAI-compatible model listing alone does not establish that support. Record the exact model weights/revision, quantization, inference-server version, served model ID, API base URL, and tool/reasoning parser configuration. Use the installed server release's documentation for parser names and launch parameters.

Confirm connectivity from PacketSafari's backend and workers, approve the host where required, and configure private CAs or proxies before testing. See egress approval and proxies and private CAs.

Connect and assign the model

Follow AI Access: provider onboarding for the existing Providers → Add connection, Models, Health, and Defaults & access controls. Installation administrators use Admin → AI settings; a team owner/admin uses My team → AI when permitted.

Under Installation AI models, add a model or customize a preset. Match the profile to the local server, including context/output limits, reasoning behavior, parser/revision details and structured-output capability. Profiles contain no credentials. JSON import/export supports reviewing a profile without discovery; import does not contact a provider or implement an unsupported API dialect. Choose permitted workflows and teams, then save.

Use Require passing validation for a route that must pass the synthetic contract before use. Enable without validation is an explicit administrator exception, not a passing test. The exact route must still be tested against representative PCAPs before production rollout. Hosted OpenRouter results do not qualify a locally served copy of the same model.

Establish a local preset

Record these alongside the acceptance results:

SettingWhat to record
IdentityEndpoint, served model ID, model/server revisions and quantization
ProtocolAPI dialect, tool parser, reasoning parser and history requirements
Request limitsActual context window and maximum output supported by this deployment
ReasoningSupported efforts, mappings if needed, and reasoning.default_effort when supported
CapacityTested active-request count, queue behavior, memory headroom and worker capacity
AccessPermitted teams, workflows, default connection/model and workload overrides
Offline policyLocal-only connections/backups and enforced network egress policy

Start with one active investigation and then test the intended simultaneous load. GPU memory alone does not determine safe context or concurrency. PacketSafari's model profile does not automatically tune inference-server admission limits. Do not send OpenAI-style reasoning or verbosity parameters unless the local route supports their declared mapping; preserve reasoning/tool history when required by the model.

For Nemotron, the shipped Medium reasoning setting describes its OpenRouter route. Treat it as a candidate starting point, not a verified local deployment recipe. Validate local parameter support, tools, report submission and evidence quality before setting it as the team's production default.

Follow the acceptance sequence and troubleshooting guide. Record a passing configuration before rollout and repeat validation after material endpoint, model, server/parser or profile changes.