Skip to main content
A curated set of open-weight foundation models is available, optimized to run on the MAX framework. For the live list of available models and their identifiers, see the Models page in the console. Models are available through two kinds of endpoints:
  • Shared endpoints — serverless, pay-as-you-go inference on a curated set of models. No deployment or setup required; call the endpoint directly with your API key.
  • Dedicated endpoints — a dedicated deployment of a model from Modular Cloud’s broader catalog, provisioned exclusively for your organization. See Create a deployment.
Check the tables below to see which models support function calling and reasoning, and which API endpoint (/v1/chat/completions or /v1/responses) each model uses.

Shared endpoints

These models run on shared, multi-tenant infrastructure and are available out of the box to every organization.

Dedicated endpoints

Custom dedicated deployments are available for these models. Contact us for more info on dedicated deployments. Model availability changes frequently as Modular adds and promotes new models. Always refer to the console for the current list.

Using model IDs

Pass the model identifier shown in the tables above as the model parameter in API requests. For example: