Back to the blog AI Infrastructure

Can You Switch AI Model Providers Later? A Portability Checklist for Self-Hosted Applications

A configurable model endpoint is only one part of portability. Use this checklist to examine prompts, tools, embeddings, evaluations, credentials and fallback plans before adopting a self-hosted AI application.

Checklist for assessing AI model provider portability in a self-hosted application

What model-provider portability means—and what it does not

AI model provider portability is the ability to move an application’s model-dependent work to another provider without an unplanned rewrite or an unacceptable change in results. It is not an all-or-nothing property: a team may be able to change a model endpoint easily while still having substantial work to migrate prompts, tools, retrieval or operational procedures.

A setting that accepts a different endpoint or API key is a useful starting point, not proof that the application is portable. The practical test is whether your important workflows still behave acceptably after the change, and whether you know what needs to be reconfigured, tested or rebuilt.

Self-hosting the application does not, by itself, make its model layer portable. Google Cloud describes AI application development as a choice that can include managed models or bringing your own open models; that choice should not be confused with a guarantee that an application can move cleanly between providers.

  • Endpoint portability: Can you change the destination and credentials through supported configuration?
  • Workflow portability: Do prompts, tool use and retrieval still work as intended?
  • Operational portability: Can your team handle data policies, access, limits and failures with the alternative provider?
  • Outcome portability: Do representative results remain acceptable under your own criteria?
What model-provider portability means—and what it does not

Map dependencies before evaluating a switch

Start with an inventory of every place the application relies on a model or provider-specific behavior. Include interactive chat, background workflows, retrieval, classification and any other model-backed task the team depends on. Do not assume that a single provider setting controls every part of the application.

For each dependency, record what it does, where its configuration lives, who owns it and how you would verify it after a change. Mark items you cannot find or cannot test; those are portability unknowns, not evidence that the feature will transfer.

  • Models: Record the model identifiers, tasks and workflow locations in use.
  • Prompts: Find system instructions, templates, formatting requirements and prompt versions.
  • Tools: List external actions, tool schemas, expected arguments and what the application should do when a tool call is unsuitable or fails.
  • Retrieval and embeddings: Identify the embedding model, preprocessing choices, vector store, index and any dependency between them.
  • Provider-specific features: Note any options or response formats the application exposes that are tied to a particular provider or model.
  • Operations: Record credentials, access controls, data-handling requirements, usage limits and the expected failure behavior.
Map dependencies before evaluating a switch

Check whether a setting is enough to change providers

Trace the configuration path from the application interface or deployment settings to each model-backed workflow. A single editable endpoint may be sufficient for one use case but not for an application that has separate settings for chat, embeddings or background tasks.

Use a staging environment or another low-risk setup for the first test. Change one dependency at a time and document every setting that must move with it. If the change requires editing application code, workflow definitions or prompt templates, include that work in your migration estimate rather than calling the switch configuration-only.

  • Can you change the endpoint, model identifier and credentials without editing application code?
  • Are settings centralized, or do separate workflows require separate changes?
  • Can the old configuration be restored quickly if the test fails?
  • Can you distinguish a provider setting from a model-specific option?
  • Does the application make it clear which provider and model handled a request?

Test prompts and tool behavior across providers

Treat prompts as part of the application, not as interchangeable text. Run the same representative inputs through the current and candidate configurations, then review the outputs against the requirements of each workflow. A prompt that is acceptable for an open-ended answer may need different scrutiny when the application expects a precise structure or a tool action.

For tool-using workflows, test the complete path: whether the application requests the intended action, whether the arguments are usable, and whether the final response accurately reflects the action’s result. Include cases where the input is incomplete, ambiguous or outside the workflow’s scope. Do not infer successful tool behavior from a convincing-looking final answer.

  • Include ordinary, edge and out-of-scope inputs from real work, with sensitive details removed where appropriate.
  • Check required output fields, formatting constraints and refusal or escalation behavior.
  • For each tool, check that the intended action and arguments are correct before allowing consequential actions.
  • Record failures and review them by workflow; an acceptable average can conceal a critical failure in one task.
  • Revise prompts only deliberately, and keep a record of the change so you can tell whether a result changed because of the provider or the prompt.

Treat embedding changes as a separate migration

If an application uses retrieval, do not treat its embedding model as merely another chat-model setting. First establish which model produced the existing vectors and how the application creates and searches its index. Then confirm that the candidate setup is compatible with the existing index and retrieval process.

If compatibility cannot be established, plan and test an index rebuild rather than assuming old vectors can be reused. Estimate the operational work: regenerating vectors, rebuilding the index, validating retrieval and deciding how the application will behave during the transition. Keep the current index available until the new retrieval path has passed your checks.

  • Record the embedding model and any text preparation or chunking choices.
  • Verify the candidate’s requirements against the existing index and vector-store setup.
  • Test whether representative queries retrieve the expected material after the change.
  • Decide whether you need a new index, how it will be built and how you will verify it.
  • Document a rollback path that returns the application to the previous model and index.

Build a small evaluation set before switching

A portability test needs a repeatable comparison, not just a few memorable examples. Assemble a small set of inputs that represents the work the application actually performs. For each item, record what a good result must contain and what would make the result unacceptable.

Run the same set through the current and candidate configurations. Review results against your criteria, including tool actions and retrieval where relevant. This is a team-specific acceptance check, not a universal benchmark: the right criteria depend on the application and the consequences of a mistake.

  • Choose examples that cover common tasks, boundary cases and known failure modes.
  • Define pass conditions before reviewing the candidate results.
  • Compare factual completeness, required structure, appropriate use of tools and retrieval relevance where applicable.
  • Track failed cases and the remediation required, not only an overall impression.
  • Keep the set and results so future provider or model changes can be compared consistently.

Plan for credentials, data, limits and failures

A technically successful connection is not the whole migration. Check the candidate provider’s terms and data-handling arrangements against your organization’s requirements, and review who can access or change credentials. Confirm that the team understands any applicable usage limits and how it will detect and respond when requests cannot be completed.

Write down the intended behavior for failed or delayed requests before relying on a fallback. A fallback should be a tested operational choice, not an assumed safety net: verify which workflows can use it, whether its outputs are acceptable, and how the application or team recognizes that a switch has occurred.

  • Confirm that the intended data can be sent to the candidate provider under your policies and agreements.
  • Identify where credentials are stored, who can access them and how they will be changed or revoked.
  • Check applicable rate or usage limits and decide how the team will monitor them.
  • Define what users see when requests fail, are delayed or cannot use a tool.
  • If you plan a fallback, test it with the evaluation set and document when and how to activate it.

Run a low-risk portability test and document the result

Choose one low-consequence workflow and test the complete migration path before changing a business-critical one. Capture the original configuration, change only what is needed, run the evaluation set and record both results and operational steps. Restore the original setup afterward unless the candidate has met your acceptance criteria.

A useful result is a short, maintainable record: what transferred unchanged, what required adjustment, what could not be verified, and what must be repeated next time. Revisit it when the application, workflow or provider configuration changes.

  • Select a workflow whose failure will not disrupt essential work.
  • Save the current configuration and define an explicit rollback step.
  • Test configuration, prompts, tools and retrieval dependencies that apply to the workflow.
  • Record evaluation results, migration effort, unresolved issues and the person responsible for each.
  • Decide whether the remaining gaps are acceptable, need mitigation or rule out the proposed switch.

Frequently asked questions

Does self-hosting an AI application make it provider-portable?

No. Self-hosting concerns where the application runs; portability depends on how its model-backed features, configuration and workflows are designed. Test the application’s actual migration path rather than assuming the hosting model settles the question.

Is an editable model endpoint proof of portability?

No. It shows that at least one connection setting may be configurable. You still need to check prompts, tools, retrieval and embeddings, evaluation results, credentials, data requirements and failure behavior.

Can I keep my existing retrieval index when changing embedding models?

Do not assume so. Check the candidate setup against the existing index and retrieval process. If compatibility is uncertain, plan a test rebuild and validate retrieval before switching production use.

How much should I test before switching providers?

Start with a small evaluation set that covers the workflows and failure cases that matter to your team. Define acceptance criteria in advance, include tools or retrieval when used, and expand the test if the consequences of a failure are significant.

Does managed hosting handle model-provider migration for me?

Not automatically. Airbip manages deployment infrastructure for applications in its catalog, including Docker workloads on Airbip cloud servers, routing and TLS automation, DNS checks, service lifecycle management and configurable backups. Those infrastructure capabilities do not establish that a given application’s model layer is portable; assess the application and its model dependencies separately.

Sources and further reading

  1. Choosing a self-hosted or managed solution for AI app development — Google Cloud
  2. Docker documentation — Docker
  3. Traefik documentation — Traefik Labs
  4. Let’s Encrypt documentation — Internet Security Research Group
  5. OWASP Top 10 for LLM Applications — OWASP