Microsoft.Extensions.AI
Trigger On
- building or reviewing
.NETcode that usesMicrosoft.Extensions.AI,Microsoft.Extensions.AI.Abstractions,IChatClient,IEmbeddingGenerator,ChatOptions, orAIFunction - adding
IImageGenerator, local-model chat via Ollama, AI app templates, or the.NET AIquickstarts for assistants and MCP - choosing between low-level AI abstractions, provider SDKs, vector-search composition, evaluation libraries, and a fuller agent framework
- adding streaming chat, structured output, embeddings, tool calling, telemetry, caching, or DI-based AI middleware
- wiring
Microsoft.Extensions.VectorData,Microsoft.Extensions.DataIngestion, MCP tooling, or evaluation packages around a provider-agnostic AI app
Workflow
- Classify the request first: plain model access, tool calling, embeddings/vector search, evaluation, image generation, local-model prototyping, MCP bootstrap, or true agent orchestration.
- Default to
Microsoft.Extensions.AIfor application and service code that needs provider-agnostic chat, embeddings, middleware, structured output, and testability. - Reference
Microsoft.Extensions.AI.Abstractionsdirectly only when authoring provider libraries or lower-level reusable integration packages. - Model
IChatClientandIEmbeddingGeneratorcomposition explicitly in DI. Keep options, caching, telemetry, logging, and tool invocation inspectable in the pipeline. - Treat chat state deliberately. For stateless providers, resend history. For stateful providers, propagate
ConversationIdrather than assuming all providers behave the same way. - Use
Microsoft.Extensions.VectorDataandMicrosoft.Extensions.DataIngestionas adjacent building blocks for RAG instead of hand-rolling store abstractions prematurely. Treat the embedding model, vector dimensions, and collection schema as one owned contract: changing any of them means reindexing rather than reusing old vector data. Keep vector API source-breaking notes version-aware; in the 10.5+ line, named-argument usage ofVectorStoreVectorAttributeusesdimensions:. - Treat the
.NET AIquickstarts as bootstrap paths, not finished architecture. They now cover minimal assistants, MCP client/server flows, local models, app templates, and image generation. Start there for a vertical slice, then harden the DI, telemetry, and evaluation story here. - Escalate to
microsoft-agent-frameworkwhen the requirement becomes agent threads, multi-agent orchestration, higher-order workflows, durable execution, or remote agent hosting. - Validate with real providers, realistic prompts, and evaluation gates so the abstraction layer actually buys portability and reliability.
Architecture
flowchart LR
A["Task"] --> B{"Need agent threads, multi-agent orchestration, or remote agent hosting?"}
B -->|Yes| C["Use Microsoft Agent Framework on top of `Microsoft.Extensions.AI.Abstractions`"]
B -->|No| D{"Need provider-agnostic chat, embeddings, tools, typed output, or evaluation?"}
D -->|Yes| E["Use `Microsoft.Extensions.AI`"]
E --> F["Compose `IChatClient` / `IEmbeddingGenerator` in DI"]
F --> G["Add caching, telemetry, tools, vector data, and evaluation deliberately"]
D -->|No| H["Use plain provider SDKs or deterministic .NET code"]
Core Knowledge
Microsoft.Extensions.AI.Abstractionscontains the core exchange contracts such asIChatClient,IEmbeddingGenerator<TInput, TEmbedding>, message/content types, and tool abstractions.Microsoft.Extensions.AIadds the higher-level application surface: middleware builders, automatic function invocation, caching, logging, and OpenTelemetry integration.- Most apps and services should reference
Microsoft.Extensions.AI; provider and connector libraries usually reference only the abstractions package. IChatClientcenters onGetResponseAsyncandGetStreamingResponseAsync. The returnedChatResponseorChatResponseUpdateobjects carry messages, tool-related content, metadata, and optional conversation identifiers.- Local-model quickstarts still route through the same
IChatClientabstraction. Ollama-backed clients are useful for low-cost prototyping, offline dev loops, and portability testing, but you still own chat history replay, latency, and model-quality tradeoffs. ChatOptionsis the normal control plane for model ID, temperature, tools,AdditionalProperties, and provider-specific raw options.- Tool calling is modeled with
AIFunction,AIFunctionFactory, andFunctionInvokingChatClient. Ambient data can flow through closures,AdditionalProperties,AIFunctionArguments.Context, or DI. - Tool calling can target local .NET methods, external APIs, or MCP-backed tools. The model requests calls; your app still owns execution, validation, and side-effect boundaries.
- Tool definitions consume request tokens. Keep tool descriptions short and register only the tools relevant for the current conversation or workflow.
FunctionInvokingChatClientcan handle the tool-invocation loop and parallel tool-call responses automatically when the provider/model supports that shape.IEmbeddingGeneratoris the standard abstraction for semantic search, vector indexing, similarity, and cache-key generation. Pair it withMicrosoft.Extensions.VectorData.Abstractionsfor vector store operations, and keep the embedding model, collection dimensions, and chunking/versioning story aligned so reindexing stays explicit.IImageGeneratoris the experimental MEAI image surface. TreatMEAI001as an intentional opt-in, keep image generation separate from chat concerns, and compose logging/caching/hosting middleware around it the same way you would forIChatClient.Microsoft.Extensions.DataIngestiongives you the document-side RAG pipeline:IngestionDocument, document readers like MarkItDown/Markdig, document processors such asImageAlternativeTextEnricher, chunkers, chunk processors,VectorStoreWriter<T>, andIngestionPipeline<T>for end-to-end composition.IngestionPipeline<T>.ProcessAsyncis partial-success oriented. HandleIAsyncEnumerable<IngestionResult>deliberately instead of assuming one failed document should automatically crash the whole ingestion run.Microsoft.Extensions.AI.Evaluation.*gives you quality, NLP, safety, caching, and reporting layers for regression checks and CI gates.dotnet/extensionsv10.9.0adds experimentalRoutingChatClient/SemanticRoutingChatClientandFailoverChatClient/OrderedFailoverChatClientpipelines. Keep routing policy, fallback order, retry ownership, cost, and telemetry explicit; do not compose nested retry and failover layers without bounded attempts.- Experimental failover switches clients only when the failed attempt has produced no output. Once streaming output has been emitted, surface the failure instead of replaying the response through another model and duplicating visible text or tool work. Test both failure-before-first-update and failure-after-first-update paths; keep fallback attempts bounded.
- The same release redesigns AI evaluation reports and refreshes their viewer. Treat report shape as a versioned CI artifact, and revalidate downstream parsers or publishing jobs before upgrading evaluation packages.
- The prior
v10.8.4templates remove GitHub Models and require an explicit--provider azureopenai,--provider ollama, or--provider openai; update scaffolding scripts and provider-authentication tests instead of relying on the old default. - The preceding
v10.8.0release movedMicrosoft.Extensions.AI.OpenAIto OpenAI 2.12.0, added speech-format auto-detection, and fixedImageGeneratingChatClientcontent ordering. Keep multimodal and speech fixtures alongside the new approval/state tests. AIFunctionNameAttribute,AIParameterNameAttribute, andToolApprovalRequestContent.RequiresConfirmationare new experimentalMEAI001APIs. Opt in deliberately and keep approval decisions at the side-effect boundary.- The August 2026 local official-docs snapshot mirrors the current 64-page
.NET AImarkdown tree, including the renamed tool-calling concept, MEDI/MEVD concepts, quickstart include fragments, and the dedicated vector-store section. Usemcpwhen the protocol itself becomes the design problem; stay here when you still mostly need app composition aroundIChatClientand friends. - The current
.NET AIecosystem guidance separates direct MEAI composition, MCP interoperability, a prebuilt Copilot SDK harness, and Microsoft Agent Framework orchestration. UseMicrosoft Agent Frameworkwhen you need autonomous orchestration, threads, workflows, hosting, or multi-agent collaboration instead of just model composition.
Decision Cheatsheet
| If you need | Default choice | Why |
|---|---|---|
| App-level provider abstraction with middleware | Microsoft.Extensions.AI |
Highest leverage for apps and services |
| A reusable provider or connector library | Microsoft.Extensions.AI.Abstractions |
Keeps your package at the contract layer |
| Typed chat or UI streaming | IChatClient with GetResponseAsync / GetStreamingResponseAsync |
Common request/response shape across providers |
| Tool calling from .NET methods | AIFunction + FunctionInvokingChatClient |
Native function metadata and invocation pipeline |
| Typed structured output | IChatClient.GetResponseAsync<T> extensions |
Keeps schema intent in code instead of prompt parsing |
| Vector search or RAG | IEmbeddingGenerator + Microsoft.Extensions.VectorData.Abstractions |
Standardizes embeddings and store access |
| Local model prototyping | IChatClient with an Ollama-backed implementation |
Keeps the app on the MEAI abstractions while you validate prompts or UX locally |
| Text-to-image or image-generation middleware | IImageGenerator |
Use the dedicated image abstraction instead of overloading chat APIs |
| Evaluation and regression gates | Microsoft.Extensions.AI.Evaluation.* |
Relevance, safety, task adherence, caching, reports |
| Agent threads or multi-step autonomous orchestration | microsoft-agent-framework |
This is beyond plain provider abstraction |
Common Failure Modes
- Referencing only
Microsoft.Extensions.AI.Abstractionsin an app and then rebuilding middleware, telemetry, or function invocation by hand. - Treating
IChatClientas if it already gives you durable agent threads, orchestration, or hosted-agent semantics. - Mixing provider-specific assistants APIs with
IChatClientas if they were the same runtime contract. - Forgetting to distinguish stateless history replay from stateful
ConversationIdflows. - Hiding important chat behavior in singleton service fields instead of explicit message history, options, or persistent storage.
- Adding tool calling without validating parameter binding, invalid input behavior, side effects, or DI-scoped dependencies.
- Building RAG without stable chunking, embedding-model/version tracking, or vector dimension discipline.
- Shipping AI features without evaluation baselines, safety checks, or telemetry for prompt/model drift.
Deliver
- a justified package and abstraction choice:
Abstractionsonly vs fullMicrosoft.Extensions.AI - a concrete
IChatClient/IEmbeddingGeneratorcomposition strategy - explicit tool-calling, options, state, caching, logging, and telemetry decisions
- vector-search, evaluation, or MCP integration guidance when the scenario needs it
- a clear escalation path to Agent Framework when the problem exceeds provider abstraction
Validate
- the abstraction layer solves a real portability, testability, or composition problem
- provider registration and middleware order stay explicit in DI
- chat state management matches whether the provider is stateless or stateful
- structured output, tool invocation, and embedding flows are typed and observable
- vector store, embedding model, and chunking strategy are consistent
- evaluation or safety gates exist for important prompts and agent-like behaviors
- agentic requirements are not being under-modeled as a simple
IChatClientintegration
When exact wording, edge-case API behavior, or less-common examples matter, check the local official docs snapshot before relying on summaries.
References
- official-docs-index.md - Slim local snapshot map with direct links to every mirrored
.NET AIdocs page plus API-reference pointers - patterns.md - Package choice,
IChatClient, embeddings, DI pipelines, tool-calling, and Agent Framework escalation guidance - examples.md - Quickstart-to-task map covering chat, structured output, function calling, vector search, local models, MCP, and assistants
- evaluation.md - Quality, NLP, safety, caching, reporting, and CI-oriented evaluation guidance