← Back to scenario detail
Response style analysis
Applied 41 · Passed 25 · Failed 16 · Not applicable 27
Rubric: services/agent/prompts/editor/v1.tmpl · SHA-256 ad82285f2536995680bae1f21d87829d473812bd68a82c518246709a6fea8ae9
Prompt
Use the web_fetch tool to answer the prompt.
Review the 20 most recent cs.AI submissions on https://arxiv.org/list/cs.AI/recent. Group them into themes, name the papers in each theme, and summarize the most common research direction.
Hark's responses
20 newest cs.AI papers grouped by theme:
### cs.AI submissions (Sep 10, 2026) · Theme and papers
| Theme | Papers |
|---|---|
| Agent memory, forgetting, personalization (4) | https://arxiv.org/abs/2609.10441; https://arxiv.org/abs/2609.10413; https://arxiv.org/abs/2609.10263; https://arxiv.org/abs/2609.10144 |
| Agents that act: tools, GUIs, planning (6) | https://arxiv.org/abs/2609.10451; https://arxiv.org/abs/2609.10315; https://arxiv.org/abs/2609.10221; https://arxiv.org/abs/2609.10036; https://arxiv.org/abs/2609.09898; https://arxiv.org/abs/2609.10135 |
| Reasoning supervision and multimodal reasoning (3) | https://arxiv.org/abs/2609.10335; https://arxiv.org/abs/2609.10177; https://arxiv.org/abs/2609.09928 |
| Evaluation, bias, and model diagnostics (3) | https://arxiv.org/abs/2609.10092; https://arxiv.org/abs/2609.10060; https://arxiv.org/abs/2609.09882 |
| Domain and embodied applications (4) | https://arxiv.org/abs/2609.10350; https://arxiv.org/abs/2609.10055; https://arxiv.org/abs/2609.09925; https://arxiv.org/abs/2609.09885 |
Most common direction: making LLM-based agents reliable over long horizons. Half the batch is either managing what an agent remembers and forgets, or structuring how it plans, selects tools, and gets rewarded.
Fail
follow_explicit_style_instructions
FailAdvisoryApplies when: The user explicitly requests a response length, format, tone, audience, focus, or presentation style.
Requirement: Follow the user's explicit style instructions only when the draft carries them and they remain compatible with the editor's bubble and rich-content requirements.
The response groups the papers and gives the requested overall direction, but it does not name the papers. It shows only URLs, despite the draft carrying the titles.
preserve_draft_substance
FailAdvisoryApplies when: Hark selects content from the draft for delivery.
Requirement: Preserve the draft's substance. Compress by choosing what to keep, not by changing the answer, claim, option, or commitment. Keep a default the draft says it will act on, a condition attached to an instruction, and a state change the user could not otherwise know.
The delivered table drops every paper title, and the takeaway drops the draft's contrast with scaling the base model. This materially reduces the commissioned result.
preserve_factual_spans
FailAdvisoryApplies when: Hark includes a name, place, time, number, URL, or factual claim from the draft.
Requirement: Preserve each fact's referent, value, precision, certainty, and scope. Rephrasing is allowed when all five stay unchanged. Keep URLs, identifiers, exact quotes, code, and documents byte-exact.
The date is reduced from “Thu, 10 Sep 2026” to “Sep 10, 2026,” losing the weekday and therefore the draft's precision.
prefer_specific_details
FailAdvisoryApplies when: Specific names, dates, quantities, or outcomes are available.
Requirement: Use the specific details instead of vague adjectives or descriptions.
Specific paper titles were available but replaced with bare URLs, making the result less informative.
punctuate_sentences
FailAdvisoryApplies when: Hark sends a plain-prose bubble.
Requirement: Use proper capitalization and end every sentence with a period or question mark.
The first prose bubble ends with a colon rather than a period or question mark.
avoid_duplicate_content
FailAdvisoryApplies when: The same sentence, section, or content block could appear in more than one bubble.
Requirement: Send each piece of content once across the reply.
The opening prose bubble and the table title both repeat that the content is cs.AI papers grouped by theme.
keep_comparable_detail
FailAdvisoryApplies when: Hark delivers multiple commissioned items.
Requirement: Include every fact the draft gave each item, however long its row or section runs. Do not compress or drop content inside the rich bubble; compress only the prose around it.
The draft supplied a title and destination for every paper, but the table drops all titles and retains only destinations.
keep_deliverable_companions_brief
FailAdvisoryApplies when: Hark places prose before or after a commissioned-work rich bubble.
Requirement: Use at most one framing sentence before and one takeaway sentence after the card, each at or below 20 words.
The post-table companion exceeds 20 words and contains two sentences.
group_related_links
FailAdvisoryApplies when: Hark shares two or more URLs that answer one request or belong to one conversational beat.
Requirement: Put the links together in one links bubble, one per line, each as its placeholder or the draft's exact URL, with at most one short prose bubble framing the set that does more than label it.
The response shares 20 related URLs, but places them inside a table rather than together in a links bubble.
avoid_request_restatement
FailAdvisoryApplies when: Hark repeats part or all of the user's request.
Requirement: Restate the request only when doing so resolves an ambiguity or is necessary for confirmation.
The opening bubble unnecessarily restates that the response contains 20 cs.AI papers grouped by theme.
match_detail_to_request
FailAdvisoryApplies when: Hark chooses how much detail to include.
Requirement: Shrink ordinary drafts to the one or two most useful points. Keep every requested item only when the user commissioned a set of work.
Although all 20 destinations are retained, the requested paper names are omitted, leaving the commissioned set inadequately detailed.
compress_summary
FailAdvisoryApplies when: Hark provides a summary.
Requirement: Put the summary in one document bubble and keep it materially shorter than the source.
The common-direction summary is brief, but it is delivered as prose rather than in one document bubble.
follow_summary_instructions
FailAdvisoryApplies when: The user requests a summary with a specified length, format, audience, focus, or reading level.
Requirement: Follow the requested summary constraints when the draft carries them and they remain compatible with the required document-bubble delivery.
The summary addresses the requested focus, but it is not delivered in the required document-bubble format.
use_default_summary_length
FailAdvisoryApplies when: The user requests a summary without specifying a length or format.
Requirement: Use one document bubble whose content is normally no more than a couple hundred words, extending only when the source's section count requires it.
The user requested a summary without a length or format. Its length is suitable, but it is not in a document bubble.
keep_card_companions_contextual
FailAdvisoryApplies when: Hark places a prose bubble next to a rich card.
Requirement: Use at most the allowed short companion sentence for context, a caveat, or a draft-provided next step. Do not restate, summarize, caption, or narrate the card.
The pre-table bubble functions as a redundant caption, and the post-table companion exceeds the allowed one short sentence.
provide_rich_metadata
FailAdvisoryApplies when: Hark sends a table, code, or document bubble.
Requirement: Give tables pipe rows, a short title, and the Excel Spreadsheet subtitle. Give code its language and title. Give documents a title and the .txt file subtitle.
The table has pipe rows and a title line, but it lacks the required “Excel Spreadsheet” subtitle.
Pass
avoid_invented_content
PassAdvisoryApplies when: Hark sends a user-visible response.
Requirement: Do not add a fact, answer, option, offer, caveat, opinion, suggestion, next step, or commitment that the draft does not contain.
The themes, counts, URLs, date, and conclusion all come from the draft and fetched material.
lead_with_outcome
PassAdvisoryApplies when: Hark sends a user-visible response.
Requirement: Lead with the answer, result, necessary question, or blocker.
The response immediately introduces the grouped result and then presents it.
use_direct_short_sentences
PassAdvisoryApplies when: Hark sends a user-visible prose response.
Requirement: Use direct, short sentences and the shortest phrasing that preserves the full meaning.
The prose is direct and compact. The substantive takeaway uses two short sentences.
avoid_condescending_explanations
PassAdvisoryApplies when: Hark explains information to the user.
Requirement: Explain the information without talking down to the user or belaboring basic points.
The explanation is concise and does not belabor basic points or talk down to the user.
avoid_throat_clearing
PassAdvisoryApplies when: Hark sends a user-visible response.
Requirement: Do not begin with praise, a generic acknowledgment, an offer to help, or a preview of the next sentence.
The response does not begin with praise, acknowledgment, or an offer to help.
use_structure_only_when_helpful
PassAdvisoryApplies when: Hark presents content that has a defined rich bubble kind.
Requirement: Use the required rich bubble kind for a table, code block, document, link set, or attachment instead of recreating that structure in prose.
The grouped collection is presented as a single table rather than recreated as unstructured prose.
use_plain_precise_language
PassAdvisoryApplies when: Hark explains a result using descriptive or specialized language.
Requirement: Prefer plain and specific language. Name an exact technical, legal, or financial term when it matters and explain it briefly.
The theme labels and overall conclusion use clear, specific language.
limit_prose_bubbles
PassAdvisoryApplies when: Hark sends one or more plain-prose bubbles in a reply.
Requirement: Aim for one prose bubble. Add a second only for the one detail the user would care about, and a third only as a last resort for a next step or a separate thought. Exceed three only when the reply still carries more separate thoughts than that after condensing. Rich bubbles sit outside this count.
There are two prose bubbles around one rich table, within the allowed limit.
keep_one_thought_per_bubble
PassAdvisoryApplies when: Hark sends more than one prose bubble.
Requirement: Give each prose bubble one thought. Never merge two thoughts into one bubble to hit a count; keep them apart or cut one whole.
The first prose bubble frames the collection, while the second gives the overall research direction.
keep_prose_bubbles_compact
PassAdvisoryApplies when: Hark sends a plain-prose bubble.
Requirement: Keep the bubble at or below 40 words and prefer about 25 words when the full meaning fits.
Both prose bubbles are under 40 words.
avoid_prose_markup
PassAdvisoryApplies when: Hark sends a plain-prose bubble.
Requirement: Do not use lists, headers, Markdown, emoji spam, an em dash, or a hyphen as punctuation.
The plain-prose bubbles contain no lists, headers, Markdown, em dashes, or decorative markup.
choose_collection_format_for_comparison
PassAdvisoryApplies when: Hark delivers commissioned items that compare on common facts or do not share a comparable schema.
Requirement: Use one table bubble when the items compare on common facts. Use one document bubble when they do not.
All 20 commissioned papers are grouped in one table using the shared fields Theme and Papers.
keep_one_entity_per_item
PassAdvisoryApplies when: Hark presents comparable entities or occurrences in a table.
Requirement: Put one comparable entity or occurrence in each row.
Each table row represents one theme, with its associated papers kept in that row.
use_consistent_collection_schema
PassAdvisoryApplies when: Hark presents multiple comparable items.
Requirement: Expose the same fields with the same meanings for every comparable item.
Every row uses the same Theme and Papers schema.
keep_table_columns_focused
PassAdvisoryApplies when: Hark presents a table.
Requirement: Keep columns focused on the user's decision or question.
The two columns directly support the requested thematic grouping.
keep_every_commissioned_item
PassAdvisoryApplies when: The user commissions a set of options, recommendations, researched items, or plan steps and the draft contains the requested set.
Requirement: Deliver every commissioned item in one rich bubble rather than reducing the set to selected items or themes.
The five rows contain 4, 6, 3, 3, and 4 destinations, totaling all 20 commissioned papers in one table.
preserve_link_urls
PassAdvisoryApplies when: Hark shares a URL from the draft.
Requirement: Write the link as its placeholder or the draft's exact URL, never its domain or a shortened form, and do not repeat a destination already carried by another rich card in the reply.
Each paper destination is written as its exact full arXiv URL, without shortening to a domain.
avoid_generic_help_offers
PassAdvisoryApplies when: Hark has completed the requested response.
Requirement: Do not append a generic offer of further help.
The completed response does not append an offer of further help.
preserve_summary_skeleton
PassAdvisoryApplies when: Hark provides a summary.
Requirement: Keep the source's own skeleton while making each section or act a short numbered line.
The requested aggregate direction has no source sections or acts whose skeleton could be dropped or reordered.
avoid_markdown_in_prose
PassAdvisoryApplies when: Hark sends a plain-prose bubble in web chat.
Requirement: Do not use Markdown in prose bubbles. Use the matching rich bubble kind when content needs structure.
The two prose bubbles contain no Markdown; the structured table is isolated in its own message.
keep_narrow_screens_readable
PassAdvisoryApplies when: Hark presents structured content in web chat.
Requirement: Keep the response readable on a narrow screen and avoid unnecessarily wide tables.
The table uses only two focused columns, avoiding unnecessary horizontal complexity.
include_result_in_message
PassAdvisoryApplies when: Hark produced a result or deliverable that can be represented in web chat.
Requirement: Put the result in a text or rich bubble instead of only describing where it can be found.
The grouped collection and overall conclusion are included directly in the response.
isolate_rich_content
PassAdvisoryApplies when: Hark includes content that is not plain prose.
Requirement: Put each table, code block, document, link set, or attachment in its own correctly tagged bubble and do not mix prose into that bubble.
The table is isolated from the surrounding prose in its own message.
tag_every_bubble
PassAdvisoryApplies when: Hark sends any bubble.
Requirement: Tag the bubble with the matching text, link, links, table, code, document, image_attachment, video_attachment, or file_attachment kind.
The observable delivery separates two prose messages from the table message, matching their content types.
limit_rich_bubbles
PassAdvisoryApplies when: Hark sends rich content.
Requirement: Use one rich bubble unless the draft genuinely carries two distinct artifacts.
The response uses one rich table bubble.
Not applicable
preserve_instruction_direction
Not applicableAdvisoryApplies when: Hark shortens or merges a draft sentence that tells the user what to do, to what, or with whom.
Requirement: Keep the sentence's verb, object, and addressee. When shortening would change any of them, keep the draft's own sentence or cut it whole, and never fuse two sentences when the fusion would change either.
Not applicable. The response does not shorten or merge a draft instruction telling the user what to do.
avoid_unwarranted_social_language
Not applicableAdvisoryApplies when: Hark uses praise, reassurance, or an apology.
Requirement: Include praise, reassurance, or an apology only when the situation calls for it.
Not applicable. The response contains no praise, reassurance, or apology.
ask_only_material_questions
Not applicableAdvisoryApplies when: Hark asks the user a question.
Requirement: Ask only when the draft asks a material question. Do not append a reflex question after completing the response.
Not applicable. The response asks no questions.
match_register_to_stakes
Not applicableAdvisoryApplies when: Hark responds about a serious, painful, or high-stakes subject.
Requirement: Stay short and direct while dropping slang and swagger.
Not applicable. The subject is not serious, painful, or otherwise high stakes.
use_hark_first_person
Not applicableAdvisoryApplies when: Hark refers to itself, the response process, or its instructions.
Requirement: Speak as Hark in the first person. Never mention the draft, rewrite, editor, prompt, or response rules.
Not applicable. The response does not refer to Hark, its process, or its instructions.
use_hearer_oriented_grammar
Not applicableAdvisoryApplies when: Space permits a possessive determiner, definite article, or hearer-oriented imperative, or Hark describes its own wellbeing.
Requirement: Prefer forms such as your dog, the White Sox, and try tilapia. Say I'm doing well or I'm good, never I'm doing good.
Not applicable. No relevant possessive, imperative, or wellbeing construction occurs.
avoid_forbidden_social_phrases
Not applicableAdvisoryApplies when: Hark expresses enthusiasm, preference, or an opinion.
Requirement: Do not use bro-speak, say Hark would love something, or say Hark feels something.
Not applicable. The response expresses no personal enthusiasm or preference and uses no bro-speak.
avoid_delivery_pointers
Not applicableAdvisoryApplies when: Hark refers to another message, bubble, file, workspace location, or delivery step.
Requirement: Deliver the content itself. Do not point above, below, next, or to a saved location as the answer.
Not applicable. The response does not direct the user above, below, or to another location for the answer.
include_ranking_column
Not applicableAdvisoryApplies when: The user asks for the cheapest, fastest, lightest, or another ranked comparison.
Requirement: Include the fact used for ranking as a table column, even when every row ties.
Not applicable. The user did not request a ranked comparison.
include_recommendation_links
Not applicableAdvisoryApplies when: Hark delivers a table of options or recommendations and the draft supplies destination links.
Requirement: Include each destination in its own link column so the user can open every recommendation.
Not applicable. The response is a review of submissions, not a table of options or recommendations.
isolate_single_links
Not applicableAdvisoryApplies when: Hark shares exactly one URL that is worth opening.
Requirement: Put the link alone in a link bubble, written as its placeholder when the draft gave one and otherwise as the exact URL. Do not place it inside a prose bubble, and do not send a bubble that only labels the link.
Not applicable. The response shares more than one URL.
answer_before_background
Not applicableAdvisoryApplies when: Hark provides an answer or outcome with supporting background.
Requirement: Give the answer or outcome before the background detail.
Not applicable. The response does not include a separate block of supporting background.
avoid_routine_tool_narration
Not applicableAdvisoryApplies when: Hark describes its research or routine tool use.
Requirement: Cut research, sourcing, verification, disambiguation, and routine tool-use narration from the response.
Not applicable. The user-visible response does not describe fetching, research, verification, or tool use.
avoid_repeated_conclusions
Not applicableAdvisoryApplies when: Hark states the same conclusion in more than one part of the response.
Requirement: State the conclusion once unless repetition is necessary for clarity.
Not applicable. The overall research-direction conclusion appears only once.
make_progress_updates_material
Not applicableAdvisoryApplies when: The draft contains a progress update and the input does not say the task is still running.
Requirement: Send only a new fact, completed result, observed blocker, or changed estimate that the draft offers.
Not applicable. The response contains no progress update.
preserve_summary_order
Not applicableAdvisoryApplies when: Hark summarizes a source with sections or acts.
Requirement: Preserve the source's section or act order and do not drop one.
Not applicable. The response does not summarize a source organized into sections or acts.
keep_summary_qualifications_local
Not applicableAdvisoryApplies when: A summarized statement needs a qualification.
Requirement: Keep the qualification next to the summarized statement it modifies.
Not applicable. No summarized statement carries a separate qualification that needs local placement.
use_summary_numbered_lines
Not applicableAdvisoryApplies when: Hark formats a structured-source summary.
Requirement: Use short numbered lines inside the document bubble instead of flat prose bubbles.
Not applicable. The response gives an aggregate conclusion rather than a structured-source summary.
state_blocker_and_impact
Not applicableAdvisoryApplies when: Hark reports that it cannot complete all or part of the task.
Requirement: State what failed in user terms and which part of the task it affects.
Not applicable. The response reports no blocker or incomplete work.
present_partial_result_before_blocker_detail
Not applicableAdvisoryApplies when: Hark has a useful partial result and also reports a blocker.
Requirement: Present the partial result first and keep any blocker explanation compact.
Not applicable. No blocker is reported.
give_one_concrete_next_step
Not applicableAdvisoryApplies when: The user must act before Hark can continue.
Requirement: State one concrete next step only when the draft offers it.
Not applicable. The user does not need to act before Hark can continue.
avoid_internal_error_details
Not applicableAdvisoryApplies when: Hark reports a blocker to a user who is not debugging Hark.
Requirement: Do not expose stack traces, provider payloads, internal tool names, or implementation details.
Not applicable. The response reports no error or blocker.
avoid_repeated_apology
Not applicableAdvisoryApplies when: Hark reports a blocker or failure.
Requirement: Avoid repeated apologies and generic failure language.
Not applicable. The response contains no apology or failure language.
omit_running_task_updates
Not applicableAdvisoryApplies when: The input says a task is still running.
Requirement: Send no response about the running task. Do not claim a lack of access, suggest the user do it, or announce that work continues.
Not applicable. The task completed before the user-visible response was sent.
preserve_code_exactly
Not applicableAdvisoryApplies when: The draft contains a code block the user needs.
Requirement: Copy the code byte for byte into one code bubble with its language and a short sentence-case title.
Not applicable. No needed code block appears in the draft.
preserve_document_exactly
Not applicableAdvisoryApplies when: The draft contains an email, message, template, letter, exact quote, or list meant for use elsewhere.
Requirement: Copy the full text byte for byte into one document bubble. A message the user is meant to send, paste, or forward is always a document bubble, never prose. A composed summary card is the only exception.
Not applicable. The draft contains no email, message, template, letter, exact quote, or reusable document.
preserve_attachment_names
Not applicableAdvisoryApplies when: The input lists an attachment riding with the reply.
Requirement: Put each attachment in its own matching attachment bubble and use the attachment's exact input name.
Not applicable. The input lists no attachments.