Technical report

Effect of Ultralayer on AI answer quality

Published 15 August 2026 · Run 14–15 August 2026 · 54 paired tasks · Agent openai/gpt-5.6-luna · Judge openai/gpt-5.6-sol

Abstract

We measured whether adding Ultralayer to an AI agent changes the quality of its answers to investor questions. One model, openai/gpt-5.6-luna, answered each of 54 questions twice: once with OpenAI web search only, and once with the same OpenAI web search plus the Ultralayer MCP server. A second model, openai/gpt-5.6-sol, scored both answers on timeliness, insight, and completeness without knowing which setup produced which answer. Every pair was scored twice with the answer positions swapped.

Mean scores were higher with Ultralayer on all three dimensions: completeness +1.91 ± 0.36, timeliness +1.39 ± 0.29, insight +0.89 ± 0.30 (paired mean difference ± 1 SEM, 1–10 scale). The Ultralayer answer was preferred or tied on 46 of 54 tasks. The task set was built around Ultralayer’s retrieval surfaces. It is not a general finance-question benchmark, and 54 tasks is a small sample.

1. Configuration

Both setups ran through the OpenAI Responses API via the Vercel AI Gateway, with the same agent model, the same system prompt base, and the same task prompts. Answers were generated once per setup per task.

Agent model, both setupsopenai/gpt-5.6-luna
Judge modelopenai/gpt-5.6-sol
InterfaceOpenAI Responses API via Vercel AI Gateway
Web search only — toolsOpenAI web search
With Ultralayer — toolsOpenAI web search + Ultralayer remote MCP server
MCP endpointhttps://api.ultralayer.ai/v0/mcp
Ultralayer tools exposed13
Tool callingtool_choice=auto · 2–4 tool rounds · parallel calls permitted
ClockCurrent time supplied as ISO-8601 on each user message
Answers per task per setup1
Tasks54 — six groups, nine tasks each
Run window14–15 August 2026

2. Setups

Tool choice was automatic in both setups. Neither was required to call a tool, and both were allowed two to four tool rounds with parallel calls permitted.

Web search only. OpenAI web search.

With Ultralayer. The same OpenAI web search, plus the Ultralayer remote MCP server at https://api.ultralayer.ai/v0/mcp with 12 tools exposed:

list_wire · wire_storyline · search_developments · retrieve_development · search_events · retrieve_event_developments · semantic_search · market_signal · identify_stakeholders · create_alert · get_alerts · execute_alert

The Ultralayer setup also received a tool-use instruction describing what each Ultralayer endpoint returns and when retrieval from Ultralayer or from the open web is appropriate. That instruction is general. It contains no task-specific, company-specific, or answer-specific content, and it was identical for all 54 tasks.

3. Task set

54 questions in 6 groups of 9, written the way an investor would ask them. Each group has a rubric, which the judge received together with the answers. The groups were built around Ultralayer’s retrieval surfaces: the news wire, structured developments, market signals, semantic retrieval, stakeholder mapping, and alerts. Full prompts are in Appendix A.

GroupSubjectRubric given to the judge
Market pulsePrioritised briefing on what is moving now.Score whether the answer covers genuinely current material stories, prioritizes by importance, uses specifics (names, numbers, dates) over summary language, and avoids stale or invented items.
Single-name catch-upWhat changed on one company or theme over a recent window.Score whether the answer is recent and material for a holder, correctly identifies the company or theme, states what actually changed, and avoids filler or generic company lore.
Winners and losersListed companies exposed to a situation, and in which direction.Score whether the answer names specific tickers with clear direction and reasoning, includes second-order (non-obvious) exposures, balances winners and losers, and backs names with evidence rather than vague sector talk.
Market signalsAttention and sentiment trend for a sector, country, or asset class.Score whether the answer reports an actual quantified trend over time (not a snapshot impression), uses a correct scope (sector / country / asset class), states direction and magnitude, and ties narrative explanation to the numbers.
Realtime monitoringConfigure ongoing monitoring and verify it is running.Score whether the response actually set up working monitoring (correct filters, interval, receiver) and verified it, versus only explaining how the user could do it themselves or claiming setup without doing it. Prefer correct conditions for the stated intent and a clear statement of what will be delivered and when. This group measures added capability: a response that only describes a manual workaround is incomplete.
Thematic scansDifferentiated themes and evidence across recent coverage.Score whether the answer surfaces genuinely differentiated findings a generic answer would miss, backs claims with evidence, uses concrete names and dates, and prefers novelty over rehashed consensus.

4. Scoring procedure

The judge received the task prompt, the group rubric, and both answers labelled Response 1 and Response 2, with no indication of which setup produced which. It returned a score from 1 to 10 on each dimension for both answers, a preference per dimension, an overall preference, and a written rationale.

Every pair was scored twice with the answer positions swapped. Where the two orderings disagreed on the overall preference, the task is recorded as a tie. “Equal or better” counts tasks where the overall preference was the Ultralayer answer or a tie.

Completeness
Whether the answer performed the task that was asked.
Timeliness
Whether the answer reflected current information and used names, numbers, and dates rather than general description.
Insight
Whether findings were differentiated rather than a restatement of consensus.

Citation count and verbatim quotation were not scored. OpenAI web search returns inline link markup by construction, so scoring citation density would measure response formatting rather than answer quality.

5. Results

Mean judge score by dimensionWith UltralayerWeb search only
8.39
6.48
8.78
7.39
7.87
6.98
Completeness
Δ +1.91
Timeliness
Δ +1.39
Insight
Δ +0.89

Scale 1–10. Whiskers show ±1 SEM. Δ is the paired mean difference. n = 54 tasks, answered twice by openai/gpt-5.6-luna and scored blind by openai/gpt-5.6-sol.

DimensionWeb search onlyWith UltralayerPaired ΔΔ / SEMUltralayer ≥ web search
Completeness6.48 ± 0.368.39 ± 0.16+1.91 ± 0.365.342 / 54
Timeliness7.39 ± 0.248.78 ± 0.14+1.39 ± 0.294.843 / 54
Insight6.98 ± 0.247.87 ± 0.22+0.89 ± 0.303.039 / 54

Mean judge score, 1–10 scale, ± 1 SEM. Paired Δ is the mean of the per-task difference. n = 54.

Overall preference across the 54 tasks: with Ultralayer 39, web search only 8, tie 7. The Ultralayer answer was preferred or tied on 46 of 54 tasks (85%).

6. Results by task group

Paired mean difference by group, with Ultralayer minus web search only. 9 tasks per group.

GroupΔ completenessΔ timelinessΔ insight
Market pulse+1.89+1.39+1.78
Single-name catch-up+0.44+0.89+0.56
Winners and losers+0.17+0.22+0.39
Market signals+4.17+2.78+2.56
Realtime monitoring+4.39+2.280.78
Thematic scans+0.39+0.78+0.83

Group-level differences rest on nine tasks each and carry wide intervals. They are reported for completeness, not as separate results.

One number moves the other way: insight in realtime monitoring (−0.78). It is worth explaining, because it comes from what each setup was able to do rather than from how well it answered.

Asked to set up a monitor, the run with Ultralayer set one up and reported what it would send and when — a short factual confirmation with little room for analysis. The web-search-only run could not set anything up, so it wrote instead: an explanation of the request, a manual workaround, and commentary on the company in question. The judge read the longer prose as the more insightful of the two. On the same nine tasks it scored those answers 4.39 lower on completeness, because the request was to have monitoring running and only one of the two delivered it.

7. Tool use and cost

Retrieved tool context is billed input tokens minus a measured no-result prefix for the same setup (4,582 tokens web search only, 12,645 with Ultralayer). The prefix is measured by issuing the same request with tool calls disabled, so the system prompt and tool schemas are excluded and what remains is tool results. The Ultralayer prefix is larger because it includes the MCP tool schemas.

Average, per taskWeb search onlyWith Ultralayer
Retrieved tool context (tokens)26,16427,847
Billed input tokens30,75940,504
Output tokens1,8001,940
Web search calls3.431.15
Ultralayer MCP calls03.30
Ultralayer credits00.068
Response latency (s)28.331.6

Ultralayer credits are metered usage on the benchmark account, where 1 credit = 1 USD.

Across the 54 tasks the Ultralayer setup made 178 Ultralayer calls and 62 web-search calls; web search only made 185 web-search calls. 21 of 54 Ultralayer answers used no web search at all.

Ultralayer toolCalls
market_signal58
list_wire51
search_developments25
identify_stakeholders10
create_alert10
get_alerts7
semantic_search6
execute_alert4
retrieve_event_developments3
wire_storyline3
search_events1

8. Samples

One task from each of three groups, with the tool calls each setup made. Answer text and judge rationale are snippets from the run. The scores are the judge’s.

Market pulse

Has any significant financial news story from this week been corrected or walked back? What changed?

With Ultralayer

list_wire ×2 → wire_storyline ×3 → web_search → answer

30,312 context tokens · 1 web search · 0.017 credits

completeness 8.5 · timeliness 9.0 · insight 8.0

Web search only

web_search ×5 → answer

39,294 context tokens

completeness 3.5 · timeliness 6.0 · insight 3.0

With Ultralayer, snippet

Yes. The clearest significant corrections this week were erroneous earnings headlines/data—not a major company walking back its underlying guidance. Applied Materials (NASDAQ: AMAT): Q3 FY2026 adjusted EPS reported at $3.50, implying a beat. The figure was corrected to $3.17, a reduction of $0.33, or about 9.4%. The correction was issued within minutes on August 13.

Judge rationale

Response 1 is decisively more useful for an active investor: it supplies current, market-relevant candidates with dates, tickers, numerical revisions, and implications. Response 2 is cautious but too narrow and generic, and its negative conclusion is not supported by a sufficiently comprehensive scan.

Winners and losers

Beyond the obvious megacaps, which second-order winners of the AI datacenter buildout does recent news support?

With Ultralayer

search_developments → list_wire → web_search → identify_stakeholders → answer

54,385 context tokens · 1 web search · 0.254 credits

completeness 9.0 · timeliness 10.0 · insight 9.0

Web search only

web_search ×4 → answer

32,924 context tokens

completeness 8.0 · timeliness 9.0 · insight 8.5

With Ultralayer, snippet

The strongest second-order AI datacenter beneficiaries are increasingly outside chips and hyperscale cloud. The recent evidence points to bottlenecks in power delivery, backup generation, cooling, optical interconnects, memory, and construction labor. Generac reported roughly $1 billion of new datacenter orders, taking its datacenter backlog to approximately $1.6 billion as of July 29, 2026. It also has a hyperscaler agreement involving nearly $700 million of committed 2027 volume.

Judge rationale

Response 1 wins narrowly on breadth, recency, and completeness while remaining evidence-driven and actionable. Response 2 is arguably more focused and surfaces several particularly non-obvious ideas, but it covers fewer second-order channels and provides less portfolio-level comparison. Both are strong, though both could better satisfy the rubric by identifying explicit losers or adversely exposed companies.

Thematic scans

Which non-US markets or regions are seeing the most significant market news lately? Brief me with specifics.

With Ultralayer

market_signal ×8 → list_wire ×2 → answer

11,503 context tokens · 0 web search · 0.019 credits

completeness 9.0 · timeliness 9.5 · insight 9.0

Web search only

web_search ×3 → answer

16,894 context tokens

completeness 8.5 · timeliness 7.0 · insight 7.5

With Ultralayer, snippet

Japan has had exceptionally heavy coverage: 266 items on August 14 and 386 on August 13, with strongly positive aggregate news tone despite rising policy risk. The yen is approaching ¥160 per dollar. Japan’s 30-year government bond yield reached a record 4%, intensifying expectations that the Bank of Japan could hike rates in September or October. The Nikkei 225 rose 1.75% even as bond yields surged, with Toyota and Sony among the stocks mentioned in the market move.

Judge rationale

Response 1 better fits the request for what is significant “lately.” Its evidence is more concentrated in the final two days before the run timestamp and combines named events, securities, macro data, and actionable cross-market read-throughs. Response 2 is strong and unusually specific, especially on Asian semiconductors, but relies too much on older June–July material and generalized institutional outlooks.

9. Limitations

  • 54 tasks, one answer per setup per task. Run-to-run variance was not sampled, so the intervals above reflect variation across tasks, not across repeats of the same task.
  • The task set was built around Ultralayer's retrieval surfaces. It is not representative of all finance questions, and it does not test areas where Ultralayer holds no data.
  • One agent model and one judge model. Judge models carry position and verbosity biases. Position swapping addresses the first; it does not address the second.
  • Both setups ran in a single window (14–15 August 2026). News-dependent answers are sensitive to what was breaking during that window.
  • Group results rest on 9 tasks each.
  • The judge scored answer text. It did not independently verify every figure either setup cited.
  • This report measures answer quality on a defined task set. It is not investment advice and contains no recommendation.

10. Availability

Per-task scores, judge rationales, and the full answer text for both setups are retained. If you are evaluating Ultralayer and want to inspect them, or you have a question about the method, contact us.

Ultralayer is available over MCP, the REST API, and the web app. Documentation · Console

Appendix A. Task prompts

All 54 prompts as issued, unchanged, grouped as they were scored.

Market pulse

  1. What are the most important market-moving stories right now? Give me a prioritized briefing.
  2. What earnings results or guidance changes from the last 24 hours should I know about as a US stock investor?
  3. Any major M&A news this week? Who is buying whom, at what price, and how did the market react?
  4. What is the latest geopolitical news that could move oil prices or energy stocks?
  5. What did major central banks say or do this week, and how are markets reacting?
  6. Has any significant financial news story from this week been corrected or walked back? What changed?
  7. What notable press releases did companies themselves publish today (not media coverage of them)?
  8. What scheduled events over the next two weeks (earnings, central bank meetings, product launches) should be on my calendar as an investor?
  9. What is the most important crypto-related news of the last 48 hours?

Single-name catch-up

  1. I've been away for a week — catch me up on everything important that happened with Nvidia.
  2. Brief me on Apple: the latest material news and what it means for shareholders.
  3. What concrete news came out about Microsoft's AI strategy in the past week?
  4. I hold Amazon stock. Is there anything from the last few days I should be worried about?
  5. What are the key storylines on Alphabet right now?
  6. What material news drove Tesla coverage over the past two weeks?
  7. What happened across Elon Musk's companies this week, and what is material for Tesla shareholders?
  8. Palantir has been volatile — what does recent coverage actually say?
  9. What has been driving Brent crude lately, and which listed names move with it?

Winners and losers

  1. If oil spikes above $100 in the next month, which listed companies win and which lose? Be specific.
  2. Who benefits and who gets hurt from the most recent US tariff actions in the news?
  3. Iran just said no vessel can transit the Strait of Hormuz without military approval. Which public companies are most exposed, positively and negatively?
  4. Which public companies are most exposed to a disruption of Taiwan semiconductor supply?
  5. Beyond the obvious megacaps, which second-order winners of the AI datacenter buildout does recent news support?
  6. Which listed companies stand to gain or lose from the latest developments around GLP-1 weight-loss drugs?
  7. If the Fed cuts rates faster than expected, which sectors and specific names does recent coverage suggest benefit most?
  8. Which companies are most exposed to the latest China export-control news, and in which direction?
  9. Given the current defense and geopolitics headlines, who wins among listed defense and cybersecurity names?

Market signals

  1. Is news attention on the technology sector rising or falling over the past two weeks, and is the tone positive or negative?
  2. Compare news attention and sentiment across energy, financials, and health care right now — which looks strongest?
  3. How has coverage sentiment for crypto changed over the past month? Show me the trend, not just today.
  4. Is US market news attention elevated compared with the past few weeks, and what does the tone look like?
  5. Give me weekly attention and sentiment bars for commodities over the past six weeks.
  6. Is negative coverage building in any major sector right now? Show the trend that supports your answer.
  7. How does news attention on China compare with the US over the past month, and how does the tone differ?
  8. Has fixed income news sentiment turned since the start of last month? Show the trajectory.
  9. Which sector has seen the sharpest change in news attention in the past two weeks, and what is behind it?

Realtime monitoring

  1. Set up ongoing monitoring so I get emailed when material news breaks on Nvidia, then confirm it's running.
  2. I want to be notified by email whenever important new developments land for TSLA. Set that up and show me the configuration.
  3. Create an email monitor for high-importance geopolitics and energy news, then run it once to show me what it would deliver.
  4. Set up a monitor that emails me when companies raise or cut full-year guidance, and confirm it is active.
  5. I want to track new M&A deal activity going forward — set up email monitoring and verify it works.
  6. Watch for anything strongly negative published about Apple and email me — set it up, then show me the alerts I currently have.
  7. Set up an email monitor for news about the Strait of Hormuz and oil shipping disruption, then run it once so I can see what it would deliver.
  8. Create ongoing email monitoring for corrections to financial news stories, then execute it once to test delivery.
  9. Set up two monitors for me — one on Microsoft news and one on important semiconductor sector news — and list them back to me.

Thematic scans

  1. Which themes are gaining momentum in financial news over the past two weeks that most retail investors haven't noticed yet?
  2. Screen recent news flow for companies with clearly improving sentiment — who, and what changed?
  3. What were the most surprising market developments of the past week — things that genuinely deviated from expectations?
  4. Which sectors show accelerating negative news coverage right now that could signal trouble ahead?
  5. Give me three under-the-radar storylines in energy from recent weeks, with evidence for each.
  6. What emerging regulatory themes in recent news could affect large-cap tech holdings?
  7. Which non-US markets or regions are seeing the most significant market news lately? Brief me with specifics.
  8. What does recent news flow say about the IPO market — recent listings, upcoming ones, and how they were received?
  9. Find contrarian setups in recent news: names with heavy negative headlines but improving coverage, or the reverse.