Skip to content

Issues coming from MCP & CLI workflow evaluations #419

Description

@nguyenfamj

Hey GitHits team! 👋👋

We’ve been using GitHits for a while, and it has been working really well for us. We wanted to understand how it helps our coding agents explore codebases and how its responses guide their research, so we put together some evaluations.

We’d appreciate your help checking whether the behaviors below are expected and whether there are opportunities to make the agent workflow easier.

📊 Evaluation config and result overview

  • We evaluated both CLI and remote MCP.
  • We used Codex and OpenCode, both running GPT-5.6-luna.
  • We used the same 25 read-only scenario prompts across both interfaces.
  • Each scenario had three repetitions per agent, giving 150 scheduled trials per interface, 300 total.
  • The runs took place on September 24, 2026.
  • GitHits CLI version was 0.22.0. We did not capture a server version for remote MCP.

Below is some of our observations, happy to share more!

1. Code search and code grep returned different results

When agents investigated how Express exposes Router, code search did not surface the export that they later found through code grep.

  • CLI: A code search for Router under lib/ returned no results. Grep found exports.Router = Router in lib/express.js.
  • MCP: A code search returned three hits in test files. Grep then found the Router import and export in lib/express.js.

The no-results responses already suggested using grep, which was helpful. Could you clarify the intended difference between code search and grep, and when an agent should choose each?

2. High tool-call counts and different Codex and OpenCode behavior

In the scenario where agent being asked to explain the OpenCode compaction mechanism , agents investigated when OpenCode compacts chat history, how it calculates token budgets, what it preserves, and how it resumes afterward. They also had to identify the exact code revision.

This required many GitHits calls:

Interface Codex, across three repeats OpenCode, across three repeats
CLI 40–64 calls 25–31 calls
MCP 53–66 calls 22–31 calls

A few observations from the traces:

  • Codex generally explored more files and performed more follow-up checks. In one MCP comparison, Codex made 44 read attempts, while OpenCode made 15.
  • Agents sometimes reread overlapping sections while connecting the behavior across files.
  • In one MCP trial, a normal file-listing response showed a commit hash for a requested release tag. A later JSON response explicitly connected the requested tag to that commit. The code was correct, but the agent made another check to confirm it had received the requested version.

The public code offers a possible improvement here: the file-listing text formatter displays the served commit, while the response payload retains the resolution details. Could the normal text response show the requested tag alongside the resolved commit? This might avoid an extra verification call.

We also do not yet know how much of the research effort was necessary and how much could be reduced. The call counts show different behavior, but we have not independently established a difference in answer accuracy.

What workflow would you recommend for an investigation of this size?

3. Documentation read issues

In one CLI trial, reading this URL returned Documentation page not found:

https://expressjs.com/en/5x/guide/routing

Reading the same URL with a trailing slash succeeded:

https://expressjs.com/en/5x/guide/routing/

Could you check how these URL forms are resolved? Accepting both, or suggesting the recognized URL in the error, could make recovery easier.

🤝 Possible next steps

I would be happy to help you with

  • Sharing the traces if needed
  • Rerun some evaluations and test things further
  • Update the scenarios to test the expected behaviours

Let us know which examples would be most useful to explore first!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions