- Pick the right kind of model — reasoning vs non-reasoning, and why one good reasoning model is more than enough nowadaysPick the right kind of model — reasoning vs non-reasoning, болон яагаад сайн reasoning model нэг л байхад өнөө үед хангалттай байдаг тухай
- Choose platforms by strength, not brand — ChatGPT, Claude, Gemini, Grok compared live on the same tasks: data quality vs agentic strengthChoose platforms by strength, not brand — ChatGPT, Claude, Gemini, Grok платформуудын ижил төстэй даалгавар дээрх гүйцэтгэлийг шууд харьцуулах нь: data quality vs agentic strength
- Write prompts that hold up — query prompts for one-off asks, system prompts for repeatable rolesWrite prompts that hold up — нэг удаагийн асуултад query prompt, давтаж ашиглах үүрэг даалгаварт system prompt ашиглах
- Speak the connective tissue — Agents, Skills, MCPs, API & CLI as platform-agnostic concepts, not vendor featuresSpeak the connective tissue — Agents, Skills, MCPs, API & CLI зэрэг нь ямар нэгэн vendor-ийн онцлог шинж биш, харин платформ хамаарахгүйгээр ойлгогдох концепцүүд юм
- Manage the context window in practice — the 50% rule from Module 1 turned into a daily habitManage the context window in practice — Module 1 дэх 50%-ийн дүрэм өдөр тутмын зуршил болсон нь
- Connect an AI to another platform yourself — at least once, on your own machine, before the session endsConnect an AI to another platform yourself — хичээл дуусахаас өмнө дор хаяж нэг удаа, өөрийнхөө компьютер дээр
Reasoning vs non-reasoning modelsReasoning болон non-reasoning model хоорондын ялгаа
Before we get into the practical skills, we need to understand the models themselves — because the skills only work when you know what kind of machine you're talking to.Практик ур чадвар руу орохын өмнө бид модель өөрсдөө хэрхэн ажилладагийг ойлгох хэрэгтэй — учир нь ямар төрлийн машин тай ярьж байгаагаа мэдэхгүйгээр эдгээр ур чадвар ажиллахгүй.
When all this began, models came in two types.Энэ бүхэн эхэлж байхад загварууд хоёр төрөлтэй байсан.
Non-reasoning models — they just spit out language. Sentences, paragraphs, generated directly, without a thinking step. Prompt in, prediction out.Non-reasoning models — тэд зүгээр л хэл яриаг гаргаж өгдөг. Бодолцох алхамгүйгээр шууд үүсгэгдсэн өгүүлбэр, догол мөрүүд. Prompt оруулна, таамаглал гарна.
Reasoning models — these have chain-of-thought thinking built in. They think before they answer. This is the Level 1 → Level 2 jump on the progression ladder from Module 2: the difference isn't more knowledge, it's a thinking step in front of the answer.Reasoning models — эдгээр нь дотроо chain-of-thought thinking-тэй байдаг. Тэд хариулахаасаа өмнө бодож тунгаадаг. Энэ бол Module 2-ын ахиц дэвшлийн шатлал дээрх Level 1 → Level 2 түвшний үсрэлт юм: ялгаа нь илүү их мэдлэгтээ биш, харин хариултын өмнө явагдах бодолцох алхамд оршино.
The clearest everyday example: the brain dumpӨдөр тутмын хамгийн тодорхой жишээ: brain dump (санаагаа шууд буулгах)
Take whatever is in your head and transcribe it raw — voice dictation, unorganized, half-sentences, corrections mid-thought. Hand that to a non-reasoning model and it can't make sense of it; it answers the surface of the text. Hand it to a reasoning model and it works through the mess first: ah, here's what they're actually trying to say, these are the details, these are the points, this is the conclusion — and then answers that. (This entire module was written exactly this way — dictated by voice, transcribed messy, and a reasoning model structured it.)Толгойд байгаа бүхнээ түүхийгээр нь буулга — дуут тэмдэглэл, цэгцгүй, дутуу өгүүлбэрүүд, бодлын дунд хийсэн засварууд. Үүнийг non-reasoning загварт өгвөл ойлгож чадахгүй; зөвхөн текстийн гадарга талд хариулна. Reasoning загварт өгвөл эхлээд энэ заваан зүйлийг цэгцэлнэ: аан, тэдний үнэхээр хэлэх гэсэн зүйл нь энэ юм байна, дэлгэрэнгүй нь энэ, гол санаанууд нь энэ, дүгнэлт нь энэ байна — тэгээд яг түүнд хариулна. (Энэ бүх модуль яг ингэж бичигдсэн — дуугаар dictate хийж, цэгцгүй буулгаад, reasoning загвараар бүтцэд оруулсан.)
One good reasoning model is more than enough nowadays. It covers most of what prompting tricks used to compensate for.Өнөө үед нэг сайн reasoning model байхад л хангалттай илүү гарна. Энэ нь өмнө нь prompt-ийн заль мэхээр нөхдөг байсан зүйлсийн ихэнхийг орлож чадна.
Platform comparison, hands-onПлатформын харьцуулалт, практик туршилт
Every provider's models are good at some things and bad at others. You don't learn which is which from marketing pages — you learn it by testing them on the same task and comparing. That's what we do in this section, live, on your machines: ChatGPT, Claude, Gemini, and Grok, side by side. Learn by comparison, not slides.Үйлчилгээ үзүүлэгч бүрийн модель зарим зүйлийг сайн хийдэг байхад зарим дээр нь тааруу байдаг. Аль нь ямар болохыг маркетингийн хуудаснаас олж мэдэхгүй — харин ижил даалгавар дээр туршиж үзээд харьцуулснаар мэднэ. Бид энэ хэсэгт үүнийг яг таны машин дээр, шууд хийх болно: ChatGPT, Claude, Gemini, болон Grok-ийг зэрэгцүүлэн харьцуулна. Слайд уншихаас илүүтэйгээр харьцуулж суралц.
How to read the models: two kinds of benchmarksМоделиудыг хэрхэн дүгнэх вэ: хоёр төрлийн benchmark
- Intelligence benchmarks — mathematics, physics, coding. How smart the model is at hard problems.Intelligence benchmarks — mathematics, physics, coding. Model хүнд асуудлуудыг хэр ухаалгаар шийдвэрлэж байгаа хэмжүүр.
- Agentic benchmarks — tool calling, multi-step work. How well the model uses things to get a job done.Agentic benchmarks — tool calling, multi-step work. Model нь ажлыг дуусгахын тулд өгөгдсөн зүйлсийг хэр сайн ашиглаж чадаж байгаа үзүүлэлт.
The two are more or less correlated — smarter models tend to be better agents — but they are not the same axis, and the difference shows up in daily work.Энэ хоёр нь бага сага хамааралтай — ухаалаг загварууд илүү сайн agent болох хандлагатай — гэхдээ тэд нэг ижил хэмжүүр биш бөгөөд ялгаа нь өдөр тутмын ажилд ажиглагддаг.
Context window matters too. The top models from OpenAI, Anthropic, and xAI now run around a million tokens of context. (What that means in practice — and how to manage it — is Section 5.)Context window matters too. OpenAI, Anthropic, xAI-ийн шилдэг model-ууд одоо ойролцоогоор сая орчим token-ийн context-тэйгээр ажиллаж байна. (Үүний бодит амьдрал дээрх утга болон хэрхэн удирдах тухайг Section 5-аас үзнэ.)
The providersҮйлчилгээ үзүүлэгчид
OpenAI — ChatGPT. Really good at coding right now. GPT-5.6 Sol on high, extra-high, or maximum effort is a very capable model for getting things done — and the desktop app is a very good harness (we'll get into harnesses later). Where it's weaker: data quality isn't the highest. One historical caveat worth knowing: previous generations were unreliable at pulling details out of SQL- and spreadsheet-type data — they'd miss a few details here and there. The current generation fixed that, so it's no longer a worry.OpenAI — ChatGPT. Одоогийн байдлаар код бичих тал дээр маш сайн. GPT-5.6 Sol model нь high, extra-high, эсвэл maximum effort тохиргоон дээр ажил гүйцэтгэхэд маш чадвартай бөгөөд desktop app нь маш сайн harness юм (бид harness-ийн талаар хоարулаа дэлгэрэнгүй үзэх болно). Сул тал нь: data quality хамгийн өндөрт тооцогдохгүй. Мэдэхэд илүүдэхгүй нэг түүхэн онцлог: өмнөх үеийнхэн SQL болон spreadsheet төрлийн өгөгдлөөс нарийн ширийн зүйлийг татаж авахдаа найдвартай бус байсан — зарим жижиг деталийг орхигдуулдаг байсан. Одоогийн үеийнхэн үүнийг зассан тул одоо санаа зовох зүйлгүй болсон.
Anthropic — Claude. Very good at tool calling and agentic, long-horizon tasks — give it one goal and it works toward it without getting off track. Intelligence quality is really good: Fable 5 is currently the top model, and it's a new kind of intelligence.Anthropic — Claude. Tool calling болон agentic, long-horizon tasks дээр маш сайн ажилладаг — нэг зорилго өгөхөд төөрөхгүйгээр түүнийхээ төлөө ажилладаг. Оюун ухааны чанар маш өндөр: Fable 5 нь одоогийн байдлаар шилдэг model бөгөөд энэ бол цоо шинэ төрлийн оюун ухаан юм.
In software-engineer terms: Fable 5 is the staff engineer, GPT-5.6 Sol is the senior engineer, and Sonnet 5 is the junior engineer.Software-engineer-ийн хэллэгээр бол: Fable 5 бол staff engineer, GPT-5.6 Sol бол senior engineer, харин Sonnet 5 бол junior engineer юм.
xAI — Grok. Its edge is data: it feeds straight from X, constantly fresh — both companies sit under Elon Musk. It's very cheap for the quality, and it's getting good at coding since xAI bought Cursor, the coding IDE.xAI — Grok. Гол давуу тал нь data: X платформоос шууд, тасралтгүй шинэчлэгдсэн мэдээллээр тэжээгддэг бөгөөд хоёр компани хоёүүл Elon Musk-ийн удирдлага дор байдаг. Өгч буй чанартайгаа харьцуулахад маш хямд бөгөөд xAI нь кодинг хийх зориулалттай IDE болох Cursor-ийг худалдаж авснаас хойш програмчлалын тал дээр ихээхэн сайжирч байна.
Google — Gemini. Google's advantage is also data quality — very good. The model itself benchmarks well but isn't at the very top.Google — Gemini. Google-ийн давуу тал бол мөн л data quality — маш сайн. Model өөрөө benchmark-ууд дээр сайн үр дүнтэй ч хамгийн өндөрт биш юм.
So which do you use for what?Алийг нь юунд ашиглах вэ?
| The jobҮүрэг даалгавар | Reach forСонгож ашиглах | WhyЯагаад |
|---|---|---|
| ResearchResearch | Gemini · Grok | Their data quality wins.Тэдний датаны чанар давуу тал болдог. |
| Getting work doneGetting work done | Claude · ChatGPT | Their agentic strength wins.Тэдний agentic хүчин чадал давуу тал болдог. |
My own stack right now: Fable 5 for nearly everything, and I mix in cheaper models to lower costs. I used to run a lot through Codex when 5.5 and 5.6 could solve it; these days the work mainly goes through Fable 5.Миний одоо ашиглаж буй stack: бараг бүх зүйлд Fable 5, мөн зардлаа бууруулахын тулд илүү хямд моделуудыг хольж ашигладаг. 5.5 болон 5.6 хувилбарууд шийдэж чаддаг байх үед би ихэнх ажлаа Codex-оор явуулдаг байсан; харин одоо бол ажлуудын ихэнх нь Fable 5-аар дамжиж байна.
Prompting: query prompts vs system promptsPrompting: query prompts vs system prompts
Of course we have to talk about prompting. But honestly — elaborate, "crazy" query prompts are no longer really necessary. There is a part of prompting that remains necessary, and it's the part that powers agents: the system prompt.Prompting-ийн тухай ярих нь гарцаагүй. Гэхдээ үнэн хэрэгтээ бол урт нуршсан, "солиотой" query prompts одоо тийм ч шаардлагагүй болсон. Гэхдээ л prompting-ийн зайлшгүй байх ёстой нэг хэсэг бий бөгөөд тэр нь agent-үүдийг ажиллуулдаг гол хүч болох the system prompt юм.
System prompts — where the effort goesSystem prompt — гол анхаарал хандуулах хэсэг
The simple version: you give the AI an identity, plus the standing rules it should always follow.Энгийн хувилбар: та AI-д identity, мөн байнга дагаж мөрдөх тогтмол дүрмүүдийг өгдөг.
Its system prompt would say something like: "You are an accounting agent. You log transactions and organize them according to proper accounting principles and the applicable laws." That's the shape — who it is, and the right way to do its job.Үүний system prompt нь дараах байдалтай байж болно: "You are an accounting agent. You log transactions and organize them according to proper accounting principles and the applicable laws." Энэ бол бүтцийн хэлбэр нь юм — тэр хэн бэ, мөн ажлаа ямар аргаар зөв хийх вэ гэдэг нь.
And if there's a preferred style or a particular way you want things done — say it explicitly in the system prompt. Standing instructions belong here, not repeated in every message.Хэрэв таны баримтлах дуртай стиль эсвэл ямар нэг зүйлийг хийлгэх онцгой арга зам байвал үүнийгээ system prompt дотор explicitly буюу тодорхой зааж өгөөрэй. Тогтмол мөрдөх заавар энд л хамаарах ба мессеж бүр дээр давтан бичих шаардлагагүй.
Query prompts — mostly obsolete effortQuery prompt-ууд — ихэнхдээ шаардлагагүй болсон хүчин зүйл
Recent models are so smart that heavily structured query prompts barely matter anymore. They're really good at understanding what you're trying to do — even from raw voice transcription. Nowadays I don't write structured prompts at all: I just talk while I'm walking, say whatever's on my mind, completely unstructured, and dictate it in.Сүүлийн үеийн загварууд маш ухаалаг болсон тул нарийн бүтээгдсэн query prompt бичих нь бараг чухал биш болсон. Тэд таны юу хийхийг зорьж байгааг — бүр түүхий дуут бичлэгийн хөрвүүлэлтээс ч — маш сайн ойлгодог. Одоо би бүтээгдсэн prompt огт бичдэггүй: би зүгээр л алхаж байхдаа ярьж, толгойд орж ирсэн зүйлээ огт бүтэцгүйгээр хэлж, дуугаараа бичүүлчихдэг.
This is Section 1's takeaway coming back around: a good reasoning model covers what prompting tricks used to compensate for.Энэ бол Section 1-ийн гол дүгнэлт эргэн ирж байгаа хэрэг: сайн reasoning загвар нь урьд нь prompting-ийн арга мэхээр нөхдөг байсан зүйлсийг бүрэн гүйцэтгэж чаддаг.
The skill has moved from phrasing the question well to setting up the standing instructions well.Чадвар нь асуултыг сайн томьёолохоос зааварчилгааг тогтмол бөгөөд сайн тохируулах руу шилжсэн.
Agents, Skills, MCPs, API & CLIAgents, Skills, MCPs, API & CLI
These five concepts are platform-agnostic — they exist everywhere, whatever the vendor calls them. And we don't teach them by defining terms on slides. We teach all five through one running demo: we create an agent inside Claude and equip it piece by piece. Each piece is one of the concepts.Энэ таван ойлголт нь platform-agnostic — вендор юу гэж нэрлэхээс үл хамааран хаа сайгүй байдаг. Бид тэдгээрийг слайд дээр нэр томьёо тодорхойлох замаар заадаггүй. Бид тавууланг нь нэг тасралтгүй демо ашиглан заадаг: бид Claude дотор agent үүсгэж, түүнийг хэсэг хэсгээр нь тоноглодог. Хэсэг бүр нь ойлголтуудын нэг юм.
Agent. The thing we're building — an AI with an identity (that's the system prompt from Section 3) that can actually do work, not just chat. We create it live inside Claude.Agent. Бидний бүтээж буй зүйл — зүгээр нэг чатлаад зогсохгүй бодит ажил гүйцэтгэх чадвартай, өөрийн гэсэн онцлог шинж чанартай (Section 3-ын system prompt) AI. Бид үүнийг Claude дотор шууд үүсгэнэ.
MCP. A connection that plugs the agent into another tool or platform. In the demo, we give our agent an MCP connection to Firecrawl — now it can reach out to the web and pull pages in.MCP. Agent-ийг өөр tool эсвэл платформтой холбох холболт. Үзүүлэнгийн үеэр бид agent-даа Firecrawl-ийн MCP холболт өгөх бөгөөд ингэснээр вэб рүү хандаж хуудас татаж авах боломжтой болно.
Skill. A packaged capability you hand the agent — instructions for doing one kind of job well. In the demo, we give it a frontend-design skill, so the same agent that scrapes with Firecrawl can also produce designed output.Skill. Agent-д өгч буй багц чадвар — нэг төрлийн ажлыг сайн хийхэд зориулсан зааварчилгаа. Үзүүлэнгийн үеэр бид үүнд frontend-design skill өгөх бөгөөд ингэснээр Firecrawl-аар мэдээлэл цуглуулдаг яг тэр agent нь мөн дизайн хийгдсэн үр дүнг гаргах боломжтой болно.
API. How software talks to software — how your agent grabs data from a service directly. For the demo we keep it deliberately simple: a free API, no authentication, something like live stock-market data. The point is not the data; the point is seeing an API call happen and understanding that this is the mechanism underneath everything else.API. Програм хангамж программ хангамжтай хэрхэн харилцдаг тухай — таны agent сервисээс өгөгдлийг шууд хэрхэн татаж авдаг вэ. Энэхүү үзүүлэн (demo)-д бид үүнийг зориудаар энгийн байлгах үүднээс: ямар нэгэн баталгаажуулалтгүй, free API, тухайлбал бодит цагийн хувьцааны зах зээлийн өгөгдөл шиг зүйлийг ашиглана. Гол нь өгөгдөл биш, харин API дуудлага хэрхэн хийгдэж байгааг харж, энэ нь бусад бүх зүйлийн цаана ажилладаг үндсэн механизм гэдгийг ойлгох явдал юм.
CLI. This one we go deeper on, because it deserves it. First: what the terminal is. Everything you normally do with screen, mouse, keyboard, and display — the user interface — can also be done by typing commands to the computer directly. We run a few live: mkdir to make a directory, ls to list what's there. That's commanding a computer from a terminal interface instead of a graphical one.CLI. Энэ сэдвийг бид илүү гүнзгий судална, учир нь энэ нь үүнийг шаарддаг. Нэгдүгээрт: terminal гэж юу вэ. Дэлгэц, хулгана, гар болон дэлгэц ашиглан хийдэг ердийн бүх зүйл буюу хэрэглэгчийн интерфейсийг компьютер руу шууд команд бичих замаар хийх боломжтой. Бид хэд хэдэн командыг шууд туршина: хавтас үүсгэхэд mkdir, байгаа зүйлсийг жагсааж харахад ls. Энэ бол график интерфейс биш terminal интерфейсээс компьютер удирдах арга юм.
And here's why that matters: since you can command the computer this way, you can also command programs that expose CLI commands. Once you authenticate — log into a piece of software that has a CLI — you can now drive that software from the terminal yourself, or hand the commands to your agent.Энэ яагаад чухал вэ гэвэл: та computer-ийг энэ аргаар удирдаж чадах тул CLI командаар ажилладаг programs-уудыг мөн удирдах боломжтой болно. Нэгэнт нэвтэрч баталгаажуулсан бол — CLI бүхий програм хангамжид хандаж, тэрхүү программ хангамжийг терминалаас өөрөө удирдах, эсвэл командудыг agent-д шилжүүлэх боломжтой.
That's the bridge: the terminal is how agents operate real software on your behalf.Энэ бол холбоос юм: terminal нь agent-ууд таны өмнөөс бодит програм хангамжийг ажиллуулах арга хэрэгсэл юм.
Context management — working with the AI's memoryContext management — AI-ийн санах ойд ажиллах нь
The context window is short-term memory — clearly short-term, not long-term. Module 1 gave you the concept; this is where it becomes a daily habit. Long conversations degrade; the skill is managing the window.Context window бол богино хугацааны санах ой — урт хугацааны биш, тодорхой богино хугацааных. Модуль 1 танд энэ ойлголтыг өгсөн; одоо үүнийг өдөр тутмын зуршил болгох газар. Урт яриа чанараа алддаг; гол чадвар бол window-ийг удирдах явдал юм.
You can see this right inside the agent build from Section 4 — in the terminal, or in the desktop app, it doesn't really matter, because both show the context circle: a live indicator of how full the window is. Watch that circle, and the rule is simple:Үүнийг та 4-р хэсгийн agent build дотор — terminal дээр эсвэл desktop app дээр шууд харж болно. Аль нь ч байсан ялгаагүй, учир нь хоёулаа window хэр дүүрсэнийг бодит цагт харуулдаг context circle заагчийг харуулдаг. Тэр дугуйг ажиглаарай, дүрэм нь маш энгийн:
Don't let it go past half.Талаас нь хэтрүүлж болохгүй.
Push past half, and the model used to start hallucinating — a lot. So my best practice became: whenever the context window hits half full, compact or start a new session. And because I do that consistently, I don't hit the full window anymore — honestly, I've never even tried riding it to full, because it degraded well before that. The half-full habit exists precisely so you never find out what the far end looks like.Тал хувийг нь давахаар загвар нь hallucinate хийж эхэлдэг байсан — маш их. Тиймээс миний хамгийн сайн туршлага бол: context window хагаст хүрмэгц compact хийх эсвэл шинэ session эхлүүлэх. Би үүнийг тогтмол хийдэг тул window-ийг дүүргэхээ больсон — үнэнийг хэлэхэд би дүүртэл нь ашиглахыг оролдож ч үзээгүй, учир нь түүнээс өмнө чанар нь эрс мууддаг байсан. Хагас дүүргэх зуршил нь яг цаад хязгаар нь ямар байдгийг хэзээ ч мэдэхгүй байхад зориулагдсан.
Practical habitsПрактик дадал
- The 50% rule — don't push past half the context window; at 50%, compact or start fresh (newer models stretch further, but don't ride the edge)The 50% rule — context window-ийн талаас хэтрүүлж болохгүй; 50% хүрэхэд мэдээллээ хураангуйлах эсвэл шинээр эхлэх (сүүлийн үеийн модель илүү уудам зайтай ч хязгаар дээр нь тулгаж ажиллах хэрэггүй)
- Watch the context circle — it's on screen; use itContext circle-ийг ажиглаарай — энэ нь дэлгэц дээр байна; үүнийг ашиглаарай
- Recognize when a conversation is degrading (repetition, forgotten instructions)Харилцааны чанар муудаж байгааг таних (давталт, зааварчилгааг мартах)
- Know when to start a fresh chat vs. continueХэзээ шинэ chat эхлүүлж, хэзээ үргэлжлүүлэхээ мэдэх
- Summarize before continuing ("compact" the conversation)Үргэлжлүүлэхээсээ өмнө хураангуйлах (яриаг "compact" хийх)
- Put the important context up front, not buried at message 40Чухал context-ийг 40 дэх message дээр дарахын оронд хамгийн эхэнд нь оруул