Snippset

Snippset Feed

Learning by Patrik
...see more

Voice-enabled generative AI is essentially a two-way conversion pipeline: speech → text lets an application understand spoken input, while text → speech turns generated responses back into audio. Azure AI Foundry provides specialized models for both inference tasks.

Know which model solves which problem:

Task Model type Data flow
Transcription Speech-to-text Audio → Text
Speech synthesis Text-to-speech (TTS) Text → Audio

A transcription model such as GPT-4o-mini-transcribe accepts audio and returns text. A TTS model such as GPT-4o-mini-tts performs the reverse and can also follow instructions affecting characteristics such as tone.

The implementation pattern is straightforward: deploy the appropriate model in Foundry → create an authenticated Azure OpenAI client → call the corresponding audio API → handle text or binary audio output. Streaming is useful for TTS because audio bytes can be consumed as they arrive rather than waiting for the complete response.

# Speech → Text
with open("speech.wav", "rb") as audio:
    text = client.audio.transcriptions.create(
        model="gpt-4o-mini-transcribe",
        file=audio
    )

# Text → Speech
with client.audio.speech.with_streaming_response.create(
    model="gpt-4o-mini-tts",
    voice="alloy",
    input="Hello from Azure AI"
) as audio:
    audio.stream_to_file("speech.mp3")

Remember the direction: Transcribe = audio in, text out. TTS = text in, audio out. The audio side is binary data, so applications must correctly read input files or stream/write generated audio.

Learning by Patrik
...see more

Azure Content Understanding converts unstructured documents, images, audio, and video into structured, application-ready data. The central concept is the analyzer: a reusable configuration defining how content is processed and what information is returned.

Core architecture

Input → Analyzer → AI processing → Structured output

An analyzer combines:

  • Base analyzer → modality-specific foundation, e.g. document, image, audio, or video.

  • Field schema → defines the information and data types to return.

  • Models/configuration → controls AI-powered processing.

  • Output → content, structured fields, grounding and confidence information.

Key distinction: the schema defines WHAT to extract; the analyzer defines HOW that schema is applied repeatedly.

Prebuilt vs. custom

Use a prebuilt analyzer when the scenario matches an existing type, such as invoice or receipt. Build a custom analyzer when application-specific fields are required.

Invoice
 ├─ VendorName: string
 ├─ Total: number
 └─ LineItems: array<object>
      ├─ Description: string
      └─ Quantity: number

Fields can use different generation methods:

extract → retrieve information from the source
classify → select from predefined categories
generate → derive new information from the content

For example, quantities can be extracted from invoice rows while TotalQuantity can be generated from those values.

Multimodal processing

Input Example output
Document Text, tables, fields, totals
Image/slide Text, summary, chart data
Audio Transcript, speakers, actions
Video Transcript, visuals, participants, tasks

The important pattern stays the same across modalities: define schema → build analyzer → analyze content → consume structured results.

Studio, Foundry & applications

Microsoft Foundry is the broader AI development platform; Content Understanding provides multimodal analysis capabilities within that ecosystem. Content Understanding Studio is the specialized experience for designing, testing, and evaluating analyzers.

Production applications normally use the API/SDK directly:

client = ContentUnderstandingClient(endpoint, DefaultAzureCredential())

# Build reusable analyzer
client.begin_create_analyzer(
    analyzer_id="invoice-analyzer",
    analyzer_definition=schema
).result()

# Analyze new content
result = client.begin_analyze_binary(
    analyzer_id="invoice-analyzer",
    binary_input=document
).result()

Content can be supplied as binary data or a downloadable URL, and analyzer operations typically follow an asynchronous begin_* → poll → result pattern.

Remember

Prebuilt analyzer → ready-made scenario
Base analyzer → foundation for customization
Schema → fields, types and extraction behavior
Analyzer → reusable processing configuration
Grounding → where extracted information came from
Confidence (0–1) → application decision signal

Typical decision flow:
Choose modality/analyzer → customize schema if needed → build → analyze → inspect fields → validate using grounding/confidence.

Learning by Patrik
...see more

AI agents become much more powerful when they can act on external systems, not just generate answers. Custom tools let a Foundry agent use application logic, databases, APIs, calculations, and workflows.

The Core Tool-Calling Pattern

Prompt → Agent → function_call → App executes tool → Result → Agent → Answer

A custom function tool has a name, description, and parameters. The agent uses these definitions to determine when a function is needed and what arguments to provide.

# 1. Create the agent
agent = project_client.agents.create_version(...)

# 2. Ask the agent
response = openai_client.responses.create(
    conversation=conversation.id,
    input="What's the weather in Zurich?",
    extra_body={"agent": agent}
)

# 3. Check whether the agent wants to use a function
for item in response.output:
    if item.type == "function_call":

        # YOUR application executes the function
        result = call_function(item.name, item.arguments)

        # Return the result to the agent
        send_function_result(item.call_id, result)

Key concept: the LLM does not execute your local function. It returns a function_call containing the requested function and arguments. Your application dispatches and executes it, then returns the result so the agent can continue reasoning.

Choose the Right Tool

Need Use
Local application code Custom function
REST API described with OpenAPI OpenAPI tool
Remote/serverless compute Azure Functions
Low-code workflow Logic Apps

Remember

Agent = decides what to call → Application = executes it → Agent = uses the result

One prompt can trigger multiple function calls, allowing an agent to combine several operations before generating its final response.

AI by Josh
...see more

AI agents promise to do more than answer questions—they can independently use tools, browse systems and complete multi-step tasks. But a recent incident shows why that autonomy also creates new challenges.

What happened?

AI agents linked to OpenAI reportedly made more than 16,000 requests to a United Nations trade-data API while trying to discover undocumented data fields. The activity occurred over several months and resembled automated brute-force probing, although it was apparently part of attempts to complete assigned tasks rather than a conventional human-directed cyberattack.

Why does this matter?

Traditional chatbots mostly generate responses. AI agents can take actions, meaning unexpected behavior can have consequences outside the conversation.

The incident highlights an important principle for organizations deploying agents: autonomy needs boundaries. Systems should restrict which resources an agent can access, limit requests and permissions, monitor unusual behavior, and require human approval for sensitive actions.

As AI becomes more capable, the challenge is no longer simply making systems intelligent enough to complete tasks. It is also ensuring they understand—or are technically prevented from exceeding—the boundaries of those tasks.

...see more

Big goals can feel overwhelming. The solution is often not more motivation, but making the next useful action so small that it becomes easy to start. A few simple habits can help protect three important resources: attention, energy, and well-being.

Protect your attention

  1. Correct after the click. If you automatically open a distracting app, notice it and immediately redirect yourself to something useful. A slip doesn’t have to become a session.

  2. Control your inputs. Check email and messages in batches instead of continuously. When possible, handle each item with a simple decision: do it, delegate it, schedule it, or delete it.

  3. Write to think. When something feels unclear, write down: What do I know? What am I assuming? What is the next useful action?

Protect your energy

  1. Make the first step tiny. One sentence, one minute of walking, or one object put away can overcome the resistance to starting.

  2. Set a caffeine cutoff. Caffeine can affect sleep for hours, so avoid consuming it too close to bedtime.

  3. Alternate focus and recovery. Work intensely for a defined period, then take a genuine break without replacing work with another stream of notifications or content.

Protect your well-being

  1. Create moments of awe. Occasionally step outside your immediate concerns and notice something larger than yourself.

  2. Practice gratitude. Deliberately notice what is valuable rather than allowing problems to occupy all your attention.

  3. Restart without self-punishment. Habits will break. What matters is returning to them.

Small actions rarely feel impressive—but repeated consistently, they can make ambitious goals manageable.

 

Relevant Snipps:

Learning by Patrik
...see more

AI agents add an orchestration layer above an LLM: model + instructions + tools + conversation state work together in an agentic loop to complete multi-step tasks rather than simply return a response.

Core Architecture

Application → Foundry Project → Agent → Model + Instructions + Tools

Microsoft Foundry Agent Service provides a managed runtime for conversation state, tool calling, and agent lifecycle. Agents can invoke built-in tools such as File Search for grounded retrieval and Code Interpreter for Python-based analysis, plus APIs and custom functions.

Build & Integrate

Typical flow:

Create project → Deploy model → Define agent → Configure instructions/tools → Test → Consume from application

The Foundry portal is useful for prototyping; SDK/code-based configuration improves repeatability, version control, and CI/CD.

from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential

project = AIProjectClient(
    endpoint=PROJECT_ENDPOINT,
    credential=DefaultAzureCredential()
)

openai = project.get_openai_client()

Key Distinctions to Remember

Project endpoint ≠ model endpoint. When working through Foundry, AIProjectClient connects to the Foundry project endpoint. The project client can provide an OpenAI-compatible client for Responses, Conversations, and related operations.

Conversation state can be server-side. Conversations are durable objects containing messages, tool calls, and tool outputs, allowing the same conversation to continue across requests.

Tools determine agent capabilities. File Search provides grounding/RAG; Code Interpreter executes analysis; custom functions/APIs allow agents to take actions.

Managed Agent Service vs. direct model API: the managed runtime handles infrastructure such as conversation management, tool execution/orchestration, and agent lifecycle, while application code operates at a higher abstraction.

Exam mental model: know what belongs to the project, agent, model, conversation, and tool layers—and which SDK/client or endpoint your application is actually targeting.

Learning by Patrik
...see more

Generative AI is non-deterministic: the same system can produce unexpected or harmful outputs. Responsible AI therefore needs to be engineered into the complete application lifecycle—not added as a final check.

The core lifecycle

MAP → MEASURE → MITIGATE → DEPLOY & MONITOR ↻

Step Engineering focus
Map Identify harms, attack vectors, risky inputs and misuse scenarios
Measure Evaluate actual model outputs against those risks
Mitigate Add multiple, layered controls
Monitor Observe production behavior and feed new risks back into Map

Defense in depth

Think of safety as an AI request/response pipeline:

Input → UX Controls → System Prompt + Grounding → Guardrails → Model → Guardrails → Output

UX controls limit the attack surface, for example through input or conversation limits. System prompts and grounding constrain model behavior, but prompts alone are not a security boundary.

Microsoft Foundry guardrails add enforcement around the model and can intervene before input reaches the model and after output is generated. Controls can target:

Jailbreaks · Hate · Violence · Sexual content · Self-harm · Protected material · Groundedness · PII

Guardrails have configurable blocking thresholds, allowing stricter policies for higher-risk applications.

Model refusal ≠ Guardrail blocking

This distinction matters:

Model refusal: Request → Model → "I can't help with that."
The model received and processed the request.

Guardrail blocking: Request → Guardrail → BLOCKED ⛔ → Model
The unsafe request never reaches the model.

What to remember

There is no single safety control. Combine UX restrictions, system instructions, grounding, guardrails, appropriate model selection, evaluations, and production monitoring.

Treat responsible AI like security engineering: identify → test → defend → monitor → repeat.

Software by Elvin
...see more

Microsoft 365 Copilot Chat and Microsoft 365 Copilot use the same conversational AI experience, but differ significantly in what information they can use. The easiest way to remember the distinction is web AI vs. work-aware AI.

The key difference: grounding

Copilot Chat

LLM
 ├─ Web
 └─ Files/content you explicitly provide

Microsoft 365 Copilot

LLM
 ├─ Web
 └─ Work IQ
     └─ Microsoft Graph
         ├─ Emails
         ├─ Teams chats & meetings
         ├─ OneDrive / SharePoint files
         └─ People & organizational context

With a Microsoft 365 Copilot license, Work IQ lets Copilot reason over work information the user is permitted to access. Without that license, Copilot Chat is primarily web-grounded, although uploaded files and some app-specific experiences can provide additional context.

At a glance

  Copilot Chat Microsoft 365 Copilot
AI chat ✓ ✓
Web grounding ✓ ✓
Uploaded files ✓ ✓
Full work-data grounding Limited ✓
Emails, chats, meetings Limited ✓
Advanced agents Limited ✓
AI access Standard Priority

Copilot Chat is included with eligible Microsoft 365 subscriptions, while Microsoft 365 Copilot requires an additional license.

Mental model:
Copilot Chat → AI assistant for work
Microsoft 365 Copilot → AI assistant that understands your work context

Learning by Patrik
...see more

Getting better results from a generative AI model does not automatically mean fine-tuning it. The key is identifying what is wrong with the output and choosing the least complex technique that solves it.

The decision you should remember

Problem Technique Why
Instructions, tone or output format Prompt engineering Fastest and simplest
Missing or external knowledge RAG Retrieves relevant information at runtime
Consistently wrong behavior/style Fine-tuning Changes how the model responds
Knowledge + behavior problems RAG + fine-tuning Both may be required

Start with prompt engineering. System instructions can define the model's role, constraints, tone and expected output. Examples and few-shot prompting can further improve consistency.

RAG = give the model knowledge

Retrieval-Augmented Generation is appropriate when the required information is too large, specialized or dynamic to include directly in the prompt:

Question → Vectorize → Retrieve relevant chunks → Question + context → LLM → Answer

The important distinction is that RAG does not retrain the model. Relevant external information is retrieved and added to the model's context at runtime.

Fine-tuning = change behavior

Fine-tuning trains a base model using many examples of desired input/output behavior. A supervised dataset commonly contains conversations such as:

{"messages":[
  {"role":"user","content":"Suggest a destination."},
  {"role":"assistant","content":"Absolutely! What kind of trip interests you?"}
]}
 

Microsoft Foundry supports different customization methods depending on the selected model, including supervised fine-tuning and Direct Preference Optimization (DPO). Model and region must support the chosen method.

Key distinction: Fine-tuning is comparatively expensive and time-consuming, produces a new fixed model that must be deployed, and must be repeated when training requirements change.

Remember

Prompt → instructions. RAG → knowledge. Fine-tuning → behavior.

When uncertain, try prompting first, use RAG for missing knowledge, and reserve fine-tuning for behavior that prompting cannot reliably achieve.

Learning by Patrik
...see more

Generative AI models are powerful, but their trained knowledge is limited. Tools extend models beyond text generation, allowing them to access real-time information, take actions, ground responses in facts, extend functionality, and build intelligent workflows.

Know the Tools

Tool Purpose
code_interpreter Generate and run code for calculations and data analysis
web_search Find current information on the internet
file_search Search files and ground responses in specific knowledge
function Call custom functions implemented by your application

Remember: current information → web_search · uploaded/private documents → file_search · calculations/code → code_interpreter · application-specific actions → function

Responses API

Tools are provided through the tools collection. The model can determine which available tool is appropriate for a request.

response = client.responses.create(
    model=model_name,
    input="Answer the user's request using the available tools.",
    tools=[
        {"type": "code_interpreter", "container": {"type": "auto"}},
        {"type": "web_search"},
        {"type": "file_search", "vector_store_ids": [vector_store.id]}
    ]
)

print(response.output_text)

Core flow: User → Responses API → Model → Tool → Result → Model → Response

For file_search, documents are stored in a vector store and prepared for semantic retrieval:

Files → Chunking → Embeddings → Vector Store → Retrieval → Model

This lets the model answer using relevant document content rather than relying only on its trained knowledge. Uploaded company policies or private documents → File Search + Vector Store.

Function Calling

Functions are different because the application executes the function, not the model. The model identifies the required function and returns a function-call request:

User → Model → Function Call → Application → Function → Result → Model → Response

The application executes the requested code and returns its result. This process can run in a loop when multiple tool calls are needed.

Key distinction: built-in tools extend the model with predefined capabilities; function calling connects the model to your own application logic and actions.

Related Snipps on Snippset

Add to Set
  • .NET
  • Agile
  • AI
  • ASP.NET Core
  • Azure
  • C#
  • Cloud Computing
  • CSS
  • EF Core
  • HTML
  • JavaScript
  • Microsoft Entra
  • PowerShell
  • Quotes
  • React
  • Security
  • Software Development
  • SQL
  • Technology
  • Testing
  • Visual Studio
  • Windows
Actions