{"id":18126,"date":"2026-06-16T19:57:26","date_gmt":"2026-06-16T19:57:26","guid":{"rendered":"https:\/\/megagon.ai\/?post_type=faq&#038;p=18126"},"modified":"2026-06-16T19:57:28","modified_gmt":"2026-06-16T19:57:28","slug":"multi-hop-reasoning-llm-function-calling","status":"publish","type":"faq","link":"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/","title":{"rendered":"Why do LLMs fail at multi-step function calling?"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">Performance degrades as the number of required function calls increases, making multi-hop reasoning across chained tool calls unreliable. In FuncBenchGen, Megagon Labs frames tool use as traversal over a function-dependency graph and tests seven open and closed LLMs under controlled complexity. GPT-5 drops from 72.5% accuracy with 5 core nodes to 15.0% with 20 core nodes. Connected irrelevant functions that share variables with the solution path further degrade state tracking, causing failures even when individual calls are syntactically valid.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Performance degrades as the number of required function calls increases, making multi-hop reasoning across chained tool calls unreliable. In FuncBenchGen, Megagon Labs frames tool use as traversal over a function-dependency graph and tests seven open and closed LLMs under controlled complexity. GPT-5 drops from 72.5% accuracy with 5 core nodes to 15.0% with 20 core [&hellip;]<\/p>\n","protected":false},"author":4,"template":"","meta":{"footnotes":""},"question-topic":[],"class_list":["post-18126","faq","type-faq","status-publish","hentry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Why do LLMs fail at multi-step function calling? - Megagon<\/title>\n<meta name=\"description\" content=\"Multi-hop reasoning fails as function-calling chains grow longer. GPT-5 drops from 72.5% to 15.0% accuracy as the number of required steps increases in testing.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/\" \/>\n<meta property=\"og:locale\" content=\"ja_JP\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Why do LLMs fail at multi-step function calling? - Megagon\" \/>\n<meta property=\"og:description\" content=\"Multi-hop reasoning fails as function-calling chains grow longer. GPT-5 drops from 72.5% to 15.0% accuracy as the number of required steps increases in testing.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/\" \/>\n<meta property=\"og:site_name\" content=\"Megagon\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/megagonlabs\/\" \/>\n<meta property=\"article:modified_time\" content=\"2026-06-16T19:57:28+00:00\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/faq\\\/multi-hop-reasoning-llm-function-calling\\\/\",\"url\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/faq\\\/multi-hop-reasoning-llm-function-calling\\\/\",\"name\":\"Why do LLMs fail at multi-step function calling? - Megagon\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/#website\"},\"datePublished\":\"2026-06-16T19:57:26+00:00\",\"dateModified\":\"2026-06-16T19:57:28+00:00\",\"description\":\"Multi-hop reasoning fails as function-calling chains grow longer. GPT-5 drops from 72.5% to 15.0% accuracy as the number of required steps increases in testing.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/faq\\\/multi-hop-reasoning-llm-function-calling\\\/#breadcrumb\"},\"inLanguage\":\"ja-JP\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/megagon.ai\\\/jp\\\/faq\\\/multi-hop-reasoning-llm-function-calling\\\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/faq\\\/multi-hop-reasoning-llm-function-calling\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Why do LLMs fail at multi-step function calling?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/#website\",\"url\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/\",\"name\":\"Megagon Labs\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"ja-JP\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/megagon.ai\\\/jp\\\/#organization\",\"name\":\"Megagon Labs\",\"url\":\"https:\\\/\\\/megagon.ai\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"url\":\"https:\\\/\\\/megagon.ai\\\/wp-content\\\/uploads\\\/2025\\\/02\\\/Logo-Megagon-Labs.webp\",\"caption\":\"Megagon Labs\"},\"image\":{\"url\":\"https:\\\/\\\/megagon.ai\\\/wp-content\\\/uploads\\\/2025\\\/02\\\/Logo-Megagon-Labs.webp\"},\"description\":\"Megagon Labs is an AI research organization conducting research in compound AI systems, large language models, data-AI symbiosis, and human-centered AI. Megagon Labs shares its findings with the broader community through open-source tools, datasets, publications, workshops, and an invited speaker series.\",\"sameAs\":[\"https:\\\/\\\/github.com\\\/megagonlabs\",\"https:\\\/\\\/twitter.com\\\/megagonlabs\",\"https:\\\/\\\/www.linkedin.com\\\/company\\\/megagon-labs\\\/\",\"https:\\\/\\\/www.facebook.com\\\/megagonlabs\\\/\"],\"address\":{\"@type\":\"PostalAddress\",\"streetAddress\":\"444 Castro Street\",\"addressLocality\":\"Mountain View\",\"addressRegion\":\"CA\",\"postalCode\":\"94041\",\"addressCountry\":\"US\"},\"contactPoint\":{\"@type\":\"ContactPoint\",\"email\":\"contactus@megagon.ai\",\"contactType\":\"general inquiries\"},\"parentOrganization\":{\"@type\":\"Organization\",\"name\":\"Recruit Holdings\",\"url\":\"https:\\\/\\\/recruit-holdings.com\\\/en\\\/\"}}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Why do LLMs fail at multi-step function calling? - Megagon","description":"Multi-hop reasoning fails as function-calling chains grow longer. GPT-5 drops from 72.5% to 15.0% accuracy as the number of required steps increases in testing.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/","og_locale":"ja_JP","og_type":"article","og_title":"Why do LLMs fail at multi-step function calling? - Megagon","og_description":"Multi-hop reasoning fails as function-calling chains grow longer. GPT-5 drops from 72.5% to 15.0% accuracy as the number of required steps increases in testing.","og_url":"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/","og_site_name":"Megagon","article_publisher":"https:\/\/www.facebook.com\/megagonlabs\/","article_modified_time":"2026-06-16T19:57:28+00:00","twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/","url":"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/","name":"Why do LLMs fail at multi-step function calling? - Megagon","isPartOf":{"@id":"https:\/\/megagon.ai\/jp\/#website"},"datePublished":"2026-06-16T19:57:26+00:00","dateModified":"2026-06-16T19:57:28+00:00","description":"Multi-hop reasoning fails as function-calling chains grow longer. GPT-5 drops from 72.5% to 15.0% accuracy as the number of required steps increases in testing.","breadcrumb":{"@id":"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/#breadcrumb"},"inLanguage":"ja-JP","potentialAction":[{"@type":"ReadAction","target":["https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/megagon.ai\/jp\/faq\/multi-hop-reasoning-llm-function-calling\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/megagon.ai\/jp\/"},{"@type":"ListItem","position":2,"name":"Why do LLMs fail at multi-step function calling?"}]},{"@type":"WebSite","@id":"https:\/\/megagon.ai\/jp\/#website","url":"https:\/\/megagon.ai\/jp\/","name":"Megagon Labs","description":"","publisher":{"@id":"https:\/\/megagon.ai\/jp\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/megagon.ai\/jp\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"ja-JP"},{"@type":"Organization","@id":"https:\/\/megagon.ai\/jp\/#organization","name":"Megagon Labs","url":"https:\/\/megagon.ai\/","logo":{"@type":"ImageObject","url":"https:\/\/megagon.ai\/wp-content\/uploads\/2025\/02\/Logo-Megagon-Labs.webp","caption":"Megagon Labs"},"image":{"url":"https:\/\/megagon.ai\/wp-content\/uploads\/2025\/02\/Logo-Megagon-Labs.webp"},"description":"Megagon Labs is an AI research organization conducting research in compound AI systems, large language models, data-AI symbiosis, and human-centered AI. Megagon Labs shares its findings with the broader community through open-source tools, datasets, publications, workshops, and an invited speaker series.","sameAs":["https:\/\/github.com\/megagonlabs","https:\/\/twitter.com\/megagonlabs","https:\/\/www.linkedin.com\/company\/megagon-labs\/","https:\/\/www.facebook.com\/megagonlabs\/"],"address":{"@type":"PostalAddress","streetAddress":"444 Castro Street","addressLocality":"Mountain View","addressRegion":"CA","postalCode":"94041","addressCountry":"US"},"contactPoint":{"@type":"ContactPoint","email":"contactus@megagon.ai","contactType":"general inquiries"},"parentOrganization":{"@type":"Organization","name":"Recruit Holdings","url":"https:\/\/recruit-holdings.com\/en\/"}}]}},"_links":{"self":[{"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/faq\/18126","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/faq"}],"about":[{"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/types\/faq"}],"author":[{"embeddable":true,"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/users\/4"}],"version-history":[{"count":1,"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/faq\/18126\/revisions"}],"predecessor-version":[{"id":18127,"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/faq\/18126\/revisions\/18127"}],"wp:attachment":[{"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/media?parent=18126"}],"wp:term":[{"taxonomy":"question-topic","embeddable":true,"href":"https:\/\/megagon.ai\/jp\/wp-json\/wp\/v2\/question-topic?post=18126"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}