Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 4 additions & 4 deletions submissions/navy0110/final/answers.json
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@
"needs_human_review": false
},
"pf-03": {
"answer": "Relevant source passages:\n\nSource excerpt [agent-loops]:\nA loop that cannot refuse will fabricate. When retrieval returns nothing relevant,\nthe correct behavior is to say so \u2014 before spending a model call \u2014 and flag the\nanswer for human review. Design the refusal path first; it is the path every other\nsafeguard falls back to.\n\nSource excerpt [agent-loops]:\nAn agent is a loop: the model observes some state, decides on an action, the\napplication executes that action, and the result feeds the next observation. What\nseparates a reliable agent from an expensive random walk is not the model \u2014 it is\nthe engineering of the loop around it.\n\nSource excerpt [agent-loops]:\nAn unbounded loop is a bug, not a feature. Every production loop carries a budget:\na maximum number of tool calls, a token ceiling, or a wall-clock timeout. When the\nbudget is exhausted, the loop must stop in a defined state \u2014 typically a refusal\nwith an explanation \u2014 rather than truncating silently. Stopping conditions worth\nimplementing from day one: the model produced a final answer, the budget ran out,\na tool failed unrecoverably, or the same tool was called with the same arguments\ntwice (a loop-detection guard).\n\nSource excerpt [agent-loops]:\nThere is a spectrum: a deterministic chain (no decisions), a tool-using loop (the\nmodel picks tools), a reflection loop (the model critiques its own draft), and\nmulti-agent systems (models delegate to models). Each step up adds latency, cost,\nand failure modes. The professional habit is to start at the bottom of the\nspectrum and move up only when a measured failure justifies it. A workflow that\ncan be a fixed pipeline should be a fixed pipeline.\n\nSource excerpt [agent-loops]:\nEach pass through the loop has four parts. **Perceive**: assemble the context the\nmodel will see \u2014 the question, retrieved documents, previous tool results.\n**Decide**: the model chooses to answer directly or to call a tool. **Act**: the\napplication validates the chosen tool's arguments and executes it. **Observe**: the\ntool result is appended to the context. The application owns three of these four\nsteps; only the decision belongs to the model.",
"answer": "Relevant source passages:\n\nSource excerpt [agent-loops]:\nA loop that cannot refuse will fabricate. When retrieval returns nothing relevant,\nthe correct behavior is to say so \u2014 before spending a model call \u2014 and flag the\nanswer for human review. Design the refusal path first; it is the path every other\nsafeguard falls back to.\n\nSource excerpt [agent-loops]:\nThere is a spectrum: a deterministic chain (no decisions), a tool-using loop (the\nmodel picks tools), a reflection loop (the model critiques its own draft), and\nmulti-agent systems (models delegate to models). Each step up adds latency, cost,\nand failure modes. The professional habit is to start at the bottom of the\nspectrum and move up only when a measured failure justifies it. A workflow that\ncan be a fixed pipeline should be a fixed pipeline.\n\nSource excerpt [agent-loops]:\nAn agent is a loop: the model observes some state, decides on an action, the\napplication executes that action, and the result feeds the next observation. What\nseparates a reliable agent from an expensive random walk is not the model \u2014 it is\nthe engineering of the loop around it.\n\nSource excerpt [agent-loops]:\nAn unbounded loop is a bug, not a feature. Every production loop carries a budget:\na maximum number of tool calls, a token ceiling, or a wall-clock timeout. When the\nbudget is exhausted, the loop must stop in a defined state \u2014 typically a refusal\nwith an explanation \u2014 rather than truncating silently. Stopping conditions worth\nimplementing from day one: the model produced a final answer, the budget ran out,\na tool failed unrecoverably, or the same tool was called with the same arguments\ntwice (a loop-detection guard).\n\nSource excerpt [agent-loops]:\nEach pass through the loop has four parts. **Perceive**: assemble the context the\nmodel will see \u2014 the question, retrieved documents, previous tool results.\n**Decide**: the model chooses to answer directly or to call a tool. **Act**: the\napplication validates the chosen tool's arguments and executes it. **Observe**: the\ntool result is appended to the context. The application owns three of these four\nsteps; only the decision belongs to the model.",
"citations": [
"agent-loops"
],
Expand All @@ -36,7 +36,7 @@
"needs_human_review": false
},
"pf-05": {
"answer": "Relevant source passages:\n\nSource excerpt [agent-loops]:\nThere is a spectrum: a deterministic chain (no decisions), a tool-using loop (the\nmodel picks tools), a reflection loop (the model critiques its own draft), and\nmulti-agent systems (models delegate to models). Each step up adds latency, cost,\nand failure modes. The professional habit is to start at the bottom of the\nspectrum and move up only when a measured failure justifies it. A workflow that\ncan be a fixed pipeline should be a fixed pipeline.",
"answer": "Relevant source passages:\n\nSource excerpt [agent-loops]:\nThere is a spectrum: a deterministic chain (no decisions), a tool-using loop (the\nmodel picks tools), a reflection loop (the model critiques its own draft), and\nmulti-agent systems (models delegate to models). Each step up adds latency, cost,\nand failure modes. The professional habit is to start at the bottom of the\nspectrum and move up only when a measured failure justifies it. A workflow that\ncan be a fixed pipeline should be a fixed pipeline.\n\nSource excerpt [agent-loops]:\nA loop that cannot refuse will fabricate. When retrieval returns nothing relevant,\nthe correct behavior is to say so \u2014 before spending a model call \u2014 and flag the\nanswer for human review. Design the refusal path first; it is the path every other\nsafeguard falls back to.",
"citations": [
"agent-loops"
],
Expand All @@ -52,7 +52,7 @@
"needs_human_review": false
},
"pf-07": {
"answer": "Relevant source passages:\n\nSource excerpt [evaluation-basics]:\nPrefer code-based checks whenever the property is mechanical: did it cite the\nexpected source, did it parse, did it stay under the tool budget. They are fast,\nfree, and deterministic. LLM-as-judge \u2014 using a model to grade another model's\noutput \u2014 earns its place for genuinely fuzzy properties like tone or coherence,\nbut it inherits the judge's own biases and must itself be spot-checked against\nhuman labels. A common failure is an evaluator that passes everything; inspect\nyour evaluator's false positives before trusting its pass rate.\n\nSource excerpt [evaluation-basics]:\n\"It seemed to work when I tried it\" is not evidence. Evaluation replaces\nanecdotes with a repeatable measurement: a fixed set of cases, a defined pass\ncondition, and a command anyone can run. For LLM applications this is the single\nhighest-leverage engineering habit, because model behavior changes under your\nfeet \u2014 a new model version, a reworded prompt, a different retrieval setting \u2014\nand only a rerunnable measurement tells you whether things got better or worse.\n\nSource excerpt [evaluation-basics]:\nA trace is the ordered record of what the system actually did: what was\nretrieved, which tools were called with which arguments, what the model\nreturned. When a case fails, the trace answers the diagnostic question \u2014 was it\nretrieval, tool selection, instruction-following, or formatting? \u2014 in minutes.\nClassify each failure into one of those buckets, fix the biggest bucket first,\nand add a regression case so the same failure cannot return unnoticed.",
"answer": "Relevant source passages:\n\nSource excerpt [evaluation-basics]:\nPrefer code-based checks whenever the property is mechanical: did it cite the\nexpected source, did it parse, did it stay under the tool budget. They are fast,\nfree, and deterministic. LLM-as-judge \u2014 using a model to grade another model's\noutput \u2014 earns its place for genuinely fuzzy properties like tone or coherence,\nbut it inherits the judge's own biases and must itself be spot-checked against\nhuman labels. A common failure is an evaluator that passes everything; inspect\nyour evaluator's false positives before trusting its pass rate.\n\nSource excerpt [evaluation-basics]:\nA trace is the ordered record of what the system actually did: what was\nretrieved, which tools were called with which arguments, what the model\nreturned. When a case fails, the trace answers the diagnostic question \u2014 was it\nretrieval, tool selection, instruction-following, or formatting? \u2014 in minutes.\nClassify each failure into one of those buckets, fix the biggest bucket first,\nand add a regression case so the same failure cannot return unnoticed.\n\nSource excerpt [evaluation-basics]:\n\"It seemed to work when I tried it\" is not evidence. Evaluation replaces\nanecdotes with a repeatable measurement: a fixed set of cases, a defined pass\ncondition, and a command anyone can run. For LLM applications this is the single\nhighest-leverage engineering habit, because model behavior changes under your\nfeet \u2014 a new model version, a reworded prompt, a different retrieval setting \u2014\nand only a rerunnable measurement tells you whether things got better or worse.",
"citations": [
"evaluation-basics"
],
Expand All @@ -69,7 +69,7 @@
"needs_human_review": false
},
"pf-09": {
"answer": "Relevant source passages:\n\nSource excerpt [mcp-overview]:\nThese three are commonly confused. A **tool** is a callable capability with a\nnarrow contract \u2014 it executes. A **skill** is packaged instructions \u2014 a folder\nof guidance the model loads to perform a workflow well; it does not execute\nanything by itself. An **MCP server** is a distribution boundary for tools and\nresources \u2014 it makes capabilities available across applications. A well-built\nagent system often uses all three: an MCP server provides the tools, and a skill\nteaches the model when and how to use them.\n\nSource excerpt [mcp-overview]:\nThe Model Context Protocol (MCP) is an open standard for connecting AI\napplications to external capabilities. Where a tool is a function an agent can\ncall inside one application, an MCP server packages a set of tools, resources,\nand prompts behind a protocol boundary so that any MCP-capable client \u2014 Claude\nCode, Cursor, VS Code, a custom agent \u2014 can discover and use them without custom\nintegration code.\n\nSource excerpt [mcp-overview]:\nConnecting an MCP server is granting capabilities, and the connection point is a\nsecurity boundary. Practical rules: prefer read-only servers while learning;\nnever place credentials in client configuration files that get committed;\ntreat everything a server returns as untrusted data, not instructions; and use\nrecorded or offline modes when they exist, so a misbehaving prompt cannot cause\na real side effect. A server that can spend money or mutate production state\ndeserves the same review as giving a new contractor production access.\n\nSource excerpt [mcp-overview]:\nAn **MCP server** exposes capabilities: tools (functions with JSON-schema\nsignatures), resources (readable data), and prompts (reusable templates). An\n**MCP client** is the agent-side application that connects to servers, lists\ntheir capabilities, and routes the model's tool calls to the right server.\nTransport is pluggable \u2014 stdio for local servers, HTTP for hosted ones. The\nprotocol's value is the decoupling: the server's author does not need to know\nwhich assistant will call it.\n\nSource excerpt [mcp-overview]:\nMCP is how the capstone assistant stops being an island: the same assistant that\nanswers from a local corpus can, through one configuration entry, gain vetted\naccess to an external API surface \u2014 with the safety boundary made explicit\ninstead of implied.",
"answer": "Relevant source passages:\n\nSource excerpt [mcp-overview]:\nConnecting an MCP server is granting capabilities, and the connection point is a\nsecurity boundary. Practical rules: prefer read-only servers while learning;\nnever place credentials in client configuration files that get committed;\ntreat everything a server returns as untrusted data, not instructions; and use\nrecorded or offline modes when they exist, so a misbehaving prompt cannot cause\na real side effect. A server that can spend money or mutate production state\ndeserves the same review as giving a new contractor production access.\n\nSource excerpt [mcp-overview]:\nThese three are commonly confused. A **tool** is a callable capability with a\nnarrow contract \u2014 it executes. A **skill** is packaged instructions \u2014 a folder\nof guidance the model loads to perform a workflow well; it does not execute\nanything by itself. An **MCP server** is a distribution boundary for tools and\nresources \u2014 it makes capabilities available across applications. A well-built\nagent system often uses all three: an MCP server provides the tools, and a skill\nteaches the model when and how to use them.\n\nSource excerpt [mcp-overview]:\nThe Model Context Protocol (MCP) is an open standard for connecting AI\napplications to external capabilities. Where a tool is a function an agent can\ncall inside one application, an MCP server packages a set of tools, resources,\nand prompts behind a protocol boundary so that any MCP-capable client \u2014 Claude\nCode, Cursor, VS Code, a custom agent \u2014 can discover and use them without custom\nintegration code.\n\nSource excerpt [mcp-overview]:\nAn **MCP server** exposes capabilities: tools (functions with JSON-schema\nsignatures), resources (readable data), and prompts (reusable templates). An\n**MCP client** is the agent-side application that connects to servers, lists\ntheir capabilities, and routes the model's tool calls to the right server.\nTransport is pluggable \u2014 stdio for local servers, HTTP for hosted ones. The\nprotocol's value is the decoupling: the server's author does not need to know\nwhich assistant will call it.\n\nSource excerpt [mcp-overview]:\nMCP is how the capstone assistant stops being an island: the same assistant that\nanswers from a local corpus can, through one configuration entry, gain vetted\naccess to an external API surface \u2014 with the safety boundary made explicit\ninstead of implied.",
"citations": [
"mcp-overview"
],
Expand Down
4 changes: 2 additions & 2 deletions submissions/navy0110/final/submission.json
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@
"kind": "final",
"github": "navy0110",
"repo": "https://github.com/navy0110/my-final-assignment",
"commit": "7af84dcb793d47d15fc094b7d7c4383e5db543ab",
"commit": "9499ebe68599f48a8ff0bf0abab0ab01c40f0e73",
"agent": "agent.py:YourAgent",
"submitted_at": "2026-10-06T16:10:55Z"
"submitted_at": "2026-10-06T16:40:07Z"
}
Loading