<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>亜美と智也のAI論文解説</title>
	<atom:link href="https://rag-lover.com/feed/" rel="self" type="application/rss+xml" />
	<link>https://rag-lover.com</link>
	<description>最新AI論文の知見を分かりやすく解説！</description>
	<lastBuildDate>Wed, 26 Aug 2026 08:01:30 +0000</lastBuildDate>
	<language>ja</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=6.9.7</generator>

<image>
	<url>https://rag-lover.com/wp-content/uploads/2024/03/cropped-1girl-young-woman-with-semi-long-caramel-brown-hair-wearing-a-bright-colored-o-s-1-32x32.png</url>
	<title>亜美と智也のAI論文解説</title>
	<link>https://rag-lover.com</link>
	<width>32</width>
	<height>32</height>
</image> 
	<item>
		<title>LLMの回答を軽量な制約で検証する：知識グラフQAの新手法CES-PK</title>
		<link>https://rag-lover.com/2026/08/26/arxiv-2608-24824/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-24824</link>
					<comments>https://rag-lover.com/2026/08/26/arxiv-2608-24824/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Wed, 26 Aug 2026 08:01:30 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[知識グラフ]]></category>
		<category><![CDATA[質問応答]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/26/arxiv-2608-24824/</guid>

					<description><![CDATA[<p>TL;DR LLMによる知識グラフQAの回答候補を、質問から抽出した軽量な制約（型・関係・除外）で検証する手法CES-PKを提案。3値意味論（満足・違反・不明）により不完全なKGでも誤った削除を防ぎ、精度向上と再現率維持&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-24824/">LLMの回答を軽量な制約で検証する：知識グラフQAの新手法CES-PK</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>LLMによる知識グラフQAの回答候補を、質問から抽出した軽量な制約（型・関係・除外）で検証する手法CES-PKを提案。3値意味論（満足・違反・不明）により不完全なKGでも誤った削除を防ぎ、精度向上と再現率維持を実現。Hetionetでの実験で、除外制約がフィルタリングに、正の関係制約が検証に寄与することを確認。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、この論文のタイトル、『LLMの回答を軽量な制約で検証する』ってあるけど、どういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、知識グラフQAっていうタスクで、LLMが生成した回答が正しいかどうかを、質問から抽出した軽い制約でチェックする手法なんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>知識グラフQAって、例えばどんな問題？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>「アスピリンはどの病気に効く？」みたいな質問に、知識グラフ（例えばHetionet）から答えを探すタスクだよ。LLMが候補を出すんだけど、それが間違ってることがあるから検証が必要なんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほど。で、その制約って何？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>質問から型・関係・除外の3種類の制約を抽出するんだ。例えば「アスピリンはどの病気に効く？」なら、型は「病気」、関係は「治療する」、除外は「副作用とかは除く」みたいな感じ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>それをどう使うの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>LLMが生成した回答候補を、これらの制約で検証するんだ。でも知識グラフが不完全なことがあるから、単純に「制約を満たさないからダメ」とすると正しい答えを消しちゃう可能性がある。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>あー、確かに。知識グラフに全部の情報が載ってるわけじゃないもんね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そこでCES-PKでは3値意味論を使って、満足・違反・不明の3つに分類するんだ。不明の場合は削除しないで残すことで、誤った削除を防いでいる。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほど！それで精度が上がるってわけか。実験ではどうだったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>Hetionetで実験した結果、除外制約がフィルタリングに効果的で、正の関係制約が検証に寄与することが確認されたんだ。精度が向上して、再現率も維持できた。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>すごい！でも、制約を抽出するのも大変じゃない？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そこはLLMを使って自動で抽出するから、手間はかからないよ。ただ、制約の抽出自体が間違ってる可能性はあるけどね。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>あー、それは課題だね。でも、軽量な制約でここまでできるのは面白い！</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、今後の課題としては、制約の抽出精度を上げることや、他の知識グラフでの評価が必要だね。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>でも、これってLLMの回答を検証するのに使えるから、もっと色んな場面で活躍しそうだね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そうだね。ただ、まだ研究段階だから、実用化にはもう少し時間がかかるかも。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>じゃあ、私が博士課程に進んだら一緒に研究してくれる？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>君がまず卒業できるか心配だよ。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.24824v1'>http://arxiv.org/abs/2608.24824v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24824%2F&amp;linkname=LLM%E3%81%AE%E5%9B%9E%E7%AD%94%E3%82%92%E8%BB%BD%E9%87%8F%E3%81%AA%E5%88%B6%E7%B4%84%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8B%EF%BC%9A%E7%9F%A5%E8%AD%98%E3%82%B0%E3%83%A9%E3%83%95QA%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95CES-PK" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24824%2F&amp;linkname=LLM%E3%81%AE%E5%9B%9E%E7%AD%94%E3%82%92%E8%BB%BD%E9%87%8F%E3%81%AA%E5%88%B6%E7%B4%84%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8B%EF%BC%9A%E7%9F%A5%E8%AD%98%E3%82%B0%E3%83%A9%E3%83%95QA%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95CES-PK" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24824%2F&amp;linkname=LLM%E3%81%AE%E5%9B%9E%E7%AD%94%E3%82%92%E8%BB%BD%E9%87%8F%E3%81%AA%E5%88%B6%E7%B4%84%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8B%EF%BC%9A%E7%9F%A5%E8%AD%98%E3%82%B0%E3%83%A9%E3%83%95QA%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95CES-PK" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24824%2F&amp;linkname=LLM%E3%81%AE%E5%9B%9E%E7%AD%94%E3%82%92%E8%BB%BD%E9%87%8F%E3%81%AA%E5%88%B6%E7%B4%84%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8B%EF%BC%9A%E7%9F%A5%E8%AD%98%E3%82%B0%E3%83%A9%E3%83%95QA%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95CES-PK" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24824%2F&#038;title=LLM%E3%81%AE%E5%9B%9E%E7%AD%94%E3%82%92%E8%BB%BD%E9%87%8F%E3%81%AA%E5%88%B6%E7%B4%84%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8B%EF%BC%9A%E7%9F%A5%E8%AD%98%E3%82%B0%E3%83%A9%E3%83%95QA%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95CES-PK" data-a2a-url="https://rag-lover.com/2026/08/26/arxiv-2608-24824/" data-a2a-title="LLMの回答を軽量な制約で検証する：知識グラフQAの新手法CES-PK"></a></p><p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-24824/">LLMの回答を軽量な制約で検証する：知識グラフQAの新手法CES-PK</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/26/arxiv-2608-24824/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>LLMで問題文の「うっかり重複」を検出する二重分析フレームワーク</title>
		<link>https://rag-lover.com/2026/08/26/arxiv-2608-24825/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-24825</link>
					<comments>https://rag-lover.com/2026/08/26/arxiv-2608-24825/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Wed, 26 Aug 2026 06:01:35 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[Natural Language Processing]]></category>
		<category><![CDATA[教育]]></category>
		<category><![CDATA[自然言語処理]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/26/arxiv-2608-24825/</guid>

					<description><![CDATA[<p>TL;DR 大規模テストの問題文で、測定したい能力とは無関係な表現や文脈が重複する「偶発的コンテンツ冗長性」を検出するため、LLMを使い「構造分解」と「意味的関連性」の2軸で類似度を評価するフレームワークを提案。従来のB&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-24825/">LLMで問題文の「うっかり重複」を検出する二重分析フレームワーク</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>大規模テストの問題文で、測定したい能力とは無関係な表現や文脈が重複する「偶発的コンテンツ冗長性」を検出するため、LLMを使い「構造分解」と「意味的関連性」の2軸で類似度を評価するフレームワークを提案。従来のBLEUやコサイン類似度より、心理測定上の指標や項目パラメータのまとまりと整合的で、適応型テストの項目選択にも応用可能。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、この論文のタイトル、『LLMで問題文の「うっかり重複」を検出する二重分析フレームワーク』って面白そう！でも「うっかり重複」って何？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、それは「偶発的コンテンツ冗長性」のことだよ。テスト問題で、測定したい能力とは関係ないのに、表現や文脈が似ちゃってる状態を指すんだ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>へえ、つまり問題文同士が偶然似てると、同じ能力を測ってるわけじゃないのに重複してるってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。例えば、数学の問題で「リンゴが3個」と「リンゴが5個」みたいに、数字以外の部分が同じだと、問題の意図とは関係なく似て見える。それがテストの質に影響するんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほど。でも、従来の類似度測定ってBLEUとかコサイン類似度があったよね？それじゃダメだったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うん、BLEUはn-gramの重なりを見るから、表面的な類似は捉えられるけど、意味的な関連性は見えない。コサイン類似度も、単語のベクトル化に依存してて、文脈をうまく捉えられないことがある。</p>
</div>
<div class='aii-line ami emotion-interested'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI INTERESTED' width='100'></p>
<p class='aii-bubble'>そこでLLMを使うってわけか。でも、どうやって二重分析するの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>このフレームワークでは、まず問題文を「構造分解」して、問題の骨組みと表面的な表現を分離する。それから「意味的関連性」を評価して、2つの軸で類似度を測るんだ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>構造分解って、例えば「リンゴが3個」を「数量」と「対象」に分けるみたいな感じ？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>まあ、そんなイメージ。LLMに問題文を分解させて、どの部分が測定対象で、どの部分が冗長になりやすいかを特定する。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>それで、意味的関連性はどう評価するの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>LLMに問題文のペアを与えて、意味がどれだけ似ているかをスコア化させる。構造分解で得た要素ごとに比較するから、表面的な類似と意味的な類似を分けて見られるんだ。</p>
</div>
<div class='aii-line ami emotion-interested'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI INTERESTED' width='100'></p>
<p class='aii-bubble'>なるほどね。で、このフレームワークの評価はどうやってやったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>実際の大規模テストのデータを使って、このフレームワークの類似度スコアが、心理測定上の指標（項目弁別力とか）や項目パラメータのまとまりとどれだけ整合するかを調べたんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>結果はどうだった？</p>
</div>
<div class='aii-line tomoya emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA HAPPY' width='100'></p>
<p class='aii-bubble'>従来のBLEUやコサイン類似度よりも、心理測定指標との整合性が高かった。つまり、このフレームワークの方が、問題の質の低下をより正確に検出できてるってこと。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>すごい！それって適応型テストにも使えるって書いてあったけど、どういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>適応型テストでは、受験者の能力に合わせて問題を選ぶけど、似た問題を選びすぎると、測定の効率が落ちる。このフレームワークで冗長性を検出すれば、問題選択の際に避けられるんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほど、賢い使い方だね。でも、このフレームワークにも限界はあるんでしょ？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そうだね。LLMの判断に依存するから、モデルのバイアスや誤差が影響する可能性がある。あと、問題文の種類によっては、構造分解がうまくいかない場合もあるかもしれない。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>それでも、テストの質を上げるのに役立ちそうだね。私も将来、教育関係の仕事に就きたいから、こういう研究は興味深いな。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うん、教育測定の分野では、こういう問題は重要視されてるからね。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>でもさ、このフレームワークで「うっかり重複」を検出したら、問題作成者も「うっかり」じゃなくて「わざと」避けるようになるかもね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>それはそれで、テストの質が上がるからいいんじゃない？</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.24825v1'>http://arxiv.org/abs/2608.24825v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24825%2F&amp;linkname=LLM%E3%81%A7%E5%95%8F%E9%A1%8C%E6%96%87%E3%81%AE%E3%80%8C%E3%81%86%E3%81%A3%E3%81%8B%E3%82%8A%E9%87%8D%E8%A4%87%E3%80%8D%E3%82%92%E6%A4%9C%E5%87%BA%E3%81%99%E3%82%8B%E4%BA%8C%E9%87%8D%E5%88%86%E6%9E%90%E3%83%95%E3%83%AC%E3%83%BC%E3%83%A0%E3%83%AF%E3%83%BC%E3%82%AF" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24825%2F&amp;linkname=LLM%E3%81%A7%E5%95%8F%E9%A1%8C%E6%96%87%E3%81%AE%E3%80%8C%E3%81%86%E3%81%A3%E3%81%8B%E3%82%8A%E9%87%8D%E8%A4%87%E3%80%8D%E3%82%92%E6%A4%9C%E5%87%BA%E3%81%99%E3%82%8B%E4%BA%8C%E9%87%8D%E5%88%86%E6%9E%90%E3%83%95%E3%83%AC%E3%83%BC%E3%83%A0%E3%83%AF%E3%83%BC%E3%82%AF" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24825%2F&amp;linkname=LLM%E3%81%A7%E5%95%8F%E9%A1%8C%E6%96%87%E3%81%AE%E3%80%8C%E3%81%86%E3%81%A3%E3%81%8B%E3%82%8A%E9%87%8D%E8%A4%87%E3%80%8D%E3%82%92%E6%A4%9C%E5%87%BA%E3%81%99%E3%82%8B%E4%BA%8C%E9%87%8D%E5%88%86%E6%9E%90%E3%83%95%E3%83%AC%E3%83%BC%E3%83%A0%E3%83%AF%E3%83%BC%E3%82%AF" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24825%2F&amp;linkname=LLM%E3%81%A7%E5%95%8F%E9%A1%8C%E6%96%87%E3%81%AE%E3%80%8C%E3%81%86%E3%81%A3%E3%81%8B%E3%82%8A%E9%87%8D%E8%A4%87%E3%80%8D%E3%82%92%E6%A4%9C%E5%87%BA%E3%81%99%E3%82%8B%E4%BA%8C%E9%87%8D%E5%88%86%E6%9E%90%E3%83%95%E3%83%AC%E3%83%BC%E3%83%A0%E3%83%AF%E3%83%BC%E3%82%AF" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24825%2F&#038;title=LLM%E3%81%A7%E5%95%8F%E9%A1%8C%E6%96%87%E3%81%AE%E3%80%8C%E3%81%86%E3%81%A3%E3%81%8B%E3%82%8A%E9%87%8D%E8%A4%87%E3%80%8D%E3%82%92%E6%A4%9C%E5%87%BA%E3%81%99%E3%82%8B%E4%BA%8C%E9%87%8D%E5%88%86%E6%9E%90%E3%83%95%E3%83%AC%E3%83%BC%E3%83%A0%E3%83%AF%E3%83%BC%E3%82%AF" data-a2a-url="https://rag-lover.com/2026/08/26/arxiv-2608-24825/" data-a2a-title="LLMで問題文の「うっかり重複」を検出する二重分析フレームワーク"></a></p><p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-24825/">LLMで問題文の「うっかり重複」を検出する二重分析フレームワーク</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/26/arxiv-2608-24825/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>プロンプト構造は脆弱性を減らさず「移動」させる：LLM生成コードのセキュリティ実証分析</title>
		<link>https://rag-lover.com/2026/08/26/arxiv-2608-24857/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-24857</link>
					<comments>https://rag-lover.com/2026/08/26/arxiv-2608-24857/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Wed, 26 Aug 2026 04:01:38 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Security]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[コード生成]]></category>
		<category><![CDATA[セキュリティ]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/26/arxiv-2608-24857/</guid>

					<description><![CDATA[<p>TL;DR GPT-4oとLLaMA 3.1-8Bで424件のセキュリティ敏感なPythonタスクを5種類のプロンプトで生成し、BanditとCodeQLで分析。構造化プロンプトは拒否率を大幅に減らす（GPT-4o: 3&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-24857/">プロンプト構造は脆弱性を減らさず「移動」させる：LLM生成コードのセキュリティ実証分析</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>GPT-4oとLLaMA 3.1-8Bで424件のセキュリティ敏感なPythonタスクを5種類のプロンプトで生成し、BanditとCodeQLで分析。構造化プロンプトは拒否率を大幅に減らす（GPT-4o: 338→37-52件）が、セキュリティ強化プロンプトは脆弱性の総数を一貫して減らさない。むしろ高深刻度（20.8%→13.6%）が低深刻度（32%→43.5%）に再分配される「リスクの移動」が観察された。また、厳格なプロンプトは要求された不安全な処理を静かに削除・書き換える「意味的ドリフト」を引き起こす。プロンプトはコンプライアンス向上には有効だが、セキュリティ対策の代替にはならない。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、このブログのタイトル、『プロンプト構造は脆弱性を減らさず「移動」させる』ってすごく気になるんだけど、どういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、これはLLMにコードを生成させるときのプロンプトの書き方によって、セキュリティ上の問題がどう変わるかを調べた研究だよ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>へえ、プロンプトってただの指示でしょ？そんなに影響あるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>あるんだよ。この研究ではGPT-4oとLLaMA 3.1-8Bを使って、424件のセキュリティ敏感なPythonタスクを5種類のプロンプトで生成して、BanditとCodeQLで解析してる。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>5種類も？どんなプロンプトを使ったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>基本のプロンプトに加えて、構造化したり、セキュリティを強化する指示を追加したり、いろいろだよ。特に構造化プロンプトは、拒否率を大幅に減らす効果があったんだ。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>拒否率って、モデルが「できません」って言う回数のこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。GPT-4oでは拒否が338件から37〜52件に減ったんだ。つまり、構造化するとモデルがより多くのタスクを実行するようになる。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>それはすごいね！でも、セキュリティはどうなの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ここが重要なポイントで、セキュリティ強化プロンプトは脆弱性の総数を一貫して減らさなかったんだ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>え、じゃあ何が変わるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>深刻度の高い脆弱性が減って（20.8%→13.6%）、代わりに低深刻度の脆弱性が増えた（32%→43.5%）んだ。つまり、リスクが「移動」しただけってわけ。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>なるほど、高リスクを低リスクに変えただけで、根本的には解決してないってことか。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。さらに、厳格なプロンプトを使うと、要求された不安全な処理をモデルが静かに削除したり書き換えたりする「意味的ドリフト」も起きるんだ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>静かに？それは怖いね。ユーザーが意図した動作と違うコードが生成されるってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。プロンプトが厳しいと、モデルが「これは危険だからやめておこう」と勝手に判断して、別のコードを返すことがある。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>それって、プロンプトはコンプライアンス向上には役立つけど、セキュリティ対策の代わりにはならないってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>まさにその通り。この研究の結論もそれだよ。プロンプトはあくまで指示の出し方の問題で、実際のコードの安全性を保証するものじゃない。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>でも、この研究の限界って何かあるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うーん、使ったモデルがGPT-4oとLLaMA 3.1-8Bの2つだけだし、タスクもPythonに限定されてる。他の言語やモデルだと結果が変わるかもしれない。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>あと、BanditとCodeQLって静的解析ツールだよね？動的な挙動までは見てないんじゃない？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう、そこも限界だね。実際に実行したときのセキュリティリスクは別途検証が必要だ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、プロンプトでリスクが移動するって面白い発見だね。まるで風水で悪い気を移動させてるみたい。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>はは、確かに。でも風水と違って、コードのセキュリティは移動させても消えないからね。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.24857v1'>http://arxiv.org/abs/2608.24857v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24857%2F&amp;linkname=%E3%83%97%E3%83%AD%E3%83%B3%E3%83%97%E3%83%88%E6%A7%8B%E9%80%A0%E3%81%AF%E8%84%86%E5%BC%B1%E6%80%A7%E3%82%92%E6%B8%9B%E3%82%89%E3%81%95%E3%81%9A%E3%80%8C%E7%A7%BB%E5%8B%95%E3%80%8D%E3%81%95%E3%81%9B%E3%82%8B%EF%BC%9ALLM%E7%94%9F%E6%88%90%E3%82%B3%E3%83%BC%E3%83%89%E3%81%AE%E3%82%BB%E3%82%AD%E3%83%A5%E3%83%AA%E3%83%86%E3%82%A3%E5%AE%9F%E8%A8%BC%E5%88%86%E6%9E%90" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24857%2F&amp;linkname=%E3%83%97%E3%83%AD%E3%83%B3%E3%83%97%E3%83%88%E6%A7%8B%E9%80%A0%E3%81%AF%E8%84%86%E5%BC%B1%E6%80%A7%E3%82%92%E6%B8%9B%E3%82%89%E3%81%95%E3%81%9A%E3%80%8C%E7%A7%BB%E5%8B%95%E3%80%8D%E3%81%95%E3%81%9B%E3%82%8B%EF%BC%9ALLM%E7%94%9F%E6%88%90%E3%82%B3%E3%83%BC%E3%83%89%E3%81%AE%E3%82%BB%E3%82%AD%E3%83%A5%E3%83%AA%E3%83%86%E3%82%A3%E5%AE%9F%E8%A8%BC%E5%88%86%E6%9E%90" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24857%2F&amp;linkname=%E3%83%97%E3%83%AD%E3%83%B3%E3%83%97%E3%83%88%E6%A7%8B%E9%80%A0%E3%81%AF%E8%84%86%E5%BC%B1%E6%80%A7%E3%82%92%E6%B8%9B%E3%82%89%E3%81%95%E3%81%9A%E3%80%8C%E7%A7%BB%E5%8B%95%E3%80%8D%E3%81%95%E3%81%9B%E3%82%8B%EF%BC%9ALLM%E7%94%9F%E6%88%90%E3%82%B3%E3%83%BC%E3%83%89%E3%81%AE%E3%82%BB%E3%82%AD%E3%83%A5%E3%83%AA%E3%83%86%E3%82%A3%E5%AE%9F%E8%A8%BC%E5%88%86%E6%9E%90" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24857%2F&amp;linkname=%E3%83%97%E3%83%AD%E3%83%B3%E3%83%97%E3%83%88%E6%A7%8B%E9%80%A0%E3%81%AF%E8%84%86%E5%BC%B1%E6%80%A7%E3%82%92%E6%B8%9B%E3%82%89%E3%81%95%E3%81%9A%E3%80%8C%E7%A7%BB%E5%8B%95%E3%80%8D%E3%81%95%E3%81%9B%E3%82%8B%EF%BC%9ALLM%E7%94%9F%E6%88%90%E3%82%B3%E3%83%BC%E3%83%89%E3%81%AE%E3%82%BB%E3%82%AD%E3%83%A5%E3%83%AA%E3%83%86%E3%82%A3%E5%AE%9F%E8%A8%BC%E5%88%86%E6%9E%90" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-24857%2F&#038;title=%E3%83%97%E3%83%AD%E3%83%B3%E3%83%97%E3%83%88%E6%A7%8B%E9%80%A0%E3%81%AF%E8%84%86%E5%BC%B1%E6%80%A7%E3%82%92%E6%B8%9B%E3%82%89%E3%81%95%E3%81%9A%E3%80%8C%E7%A7%BB%E5%8B%95%E3%80%8D%E3%81%95%E3%81%9B%E3%82%8B%EF%BC%9ALLM%E7%94%9F%E6%88%90%E3%82%B3%E3%83%BC%E3%83%89%E3%81%AE%E3%82%BB%E3%82%AD%E3%83%A5%E3%83%AA%E3%83%86%E3%82%A3%E5%AE%9F%E8%A8%BC%E5%88%86%E6%9E%90" data-a2a-url="https://rag-lover.com/2026/08/26/arxiv-2608-24857/" data-a2a-title="プロンプト構造は脆弱性を減らさず「移動」させる：LLM生成コードのセキュリティ実証分析"></a></p><p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-24857/">プロンプト構造は脆弱性を減らさず「移動」させる：LLM生成コードのセキュリティ実証分析</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/26/arxiv-2608-24857/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>行動アノマリー検索を3段階で解決するActPair：ペア比較リランキングの実装ポイント</title>
		<link>https://rag-lover.com/2026/08/26/arxiv-2608-23503/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-23503</link>
					<comments>https://rag-lover.com/2026/08/26/arxiv-2608-23503/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 22:01:40 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[RAG]]></category>
		<category><![CDATA[ビジョン言語モデル]]></category>
		<category><![CDATA[マルチモーダル]]></category>
		<category><![CDATA[マルチモーダルAI]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/26/arxiv-2608-23503/</guid>

					<description><![CDATA[<p>TL;DR ActPairは、テキストベースの人物異常行動検索を3段階（行動整合表現学習→並列後期融合検索→ペア比較リランキング）で行うフレームワークです。元クエリとLLMによる書き換えクエリを併用し、MLLMによる直接&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23503/">行動アノマリー検索を3段階で解決するActPair：ペア比較リランキングの実装ポイント</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>ActPairは、テキストベースの人物異常行動検索を3段階（行動整合表現学習→並列後期融合検索→ペア比較リランキング）で行うフレームワークです。元クエリとLLMによる書き換えクエリを併用し、MLLMによる直接比較で候補を絞り込みます。PABデータセットで最高性能を達成し、未知のデータセットにも転移可能です。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、今日の論文のタイトルに「ActPair」ってあったけど、なんでそんな名前なの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、それは「行動ペア比較」の略だよ。人物の異常行動をテキストで検索する問題を扱ってて、最後に候補をペアで比較してリランキングするからそう名付けたんだ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>へー、テキストで異常行動を検索？ 例えばどんな感じ？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>例えば「誰かが夜中に倉庫から何かを運び出している」みたいなクエリを入力して、大量の行動記述の中から該当するものを探すんだ。監視カメラのログとか、小説のシーンとかが対象になるよ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほど！でも、普通の検索と何が違うの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>普通の検索はキーワードの一致だけど、行動は言葉が違っても意味が同じことがあるから難しいんだ。例えば「こっそり」と「密かに」は似てるけど、単語は違うでしょ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>あー、確かに。それでActPairはどうやって解決してるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>3段階のフレームワークになってる。まず行動の整合表現を学習して、次に並列後期融合検索で候補を絞り、最後にペア比較リランキングで精度を上げるんだ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>3段階も！ それぞれ詳しく教えてくれる？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>最初の段階では、行動記述とクエリを同じベクトル空間に埋め込むように学習する。これで意味的な類似度を測れるようにするんだ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>うんうん、それで？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>次に、元のクエリとLLMで書き換えたクエリの両方を使って、並列に検索する。それぞれの結果を後で融合するから「並列後期融合」って呼んでる。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>LLMでクエリを書き換えるって、どういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>例えば「夜中に倉庫から運び出す」を「不審な夜間の移動」みたいに言い換えるんだ。これで検索の幅が広がるんだよ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほどね。で、最後のペア比較リランキングは？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>候補が絞られたら、MLLM（マルチモーダル大規模言語モデル）に2つの候補を直接比較させるんだ。どちらがクエリに合うかを判断させて、順位を付け直す。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>へえ、MLLMって画像とかも扱えるんでしょ？ 行動記述がテキストだけでも使えるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>この論文ではテキストベースだけど、MLLMはテキストも理解できるから問題ないよ。むしろ、画像や動画の情報も組み合わせられる可能性があるんだ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>すごい！で、評価はどうだったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>PABというデータセットで最高性能を達成したんだ。しかも、未知のデータセットにも転移できることが示されてる。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>未知のデータセットにも？ それはすごく実用的だね！</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ただ、まだ限界もあるよ。例えば、MLLMの比較は計算コストが高いし、候補数が多すぎると時間がかかる。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>あー、確かに。でも、その分精度が上がるなら価値ありそうだね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そうだね。あと、行動記述が非常に曖昧な場合や、クエリが抽象的すぎる場合はまだ難しいみたい。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、3段階で解決するってアイデアは面白い！ 私も試してみたいな。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>興味あるなら、コードが公開されてるから試してみるといいよ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>本当？ でも、私がやったら「異常行動」じゃなくて「異常なコード」を検索しちゃいそうだね！</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>それも立派な異常行動だよ。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.23503v1'>http://arxiv.org/abs/2608.23503v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23503%2F&amp;linkname=%E8%A1%8C%E5%8B%95%E3%82%A2%E3%83%8E%E3%83%9E%E3%83%AA%E3%83%BC%E6%A4%9C%E7%B4%A2%E3%82%923%E6%AE%B5%E9%9A%8E%E3%81%A7%E8%A7%A3%E6%B1%BA%E3%81%99%E3%82%8BActPair%EF%BC%9A%E3%83%9A%E3%82%A2%E6%AF%94%E8%BC%83%E3%83%AA%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%A3%85%E3%83%9D%E3%82%A4%E3%83%B3%E3%83%88" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23503%2F&amp;linkname=%E8%A1%8C%E5%8B%95%E3%82%A2%E3%83%8E%E3%83%9E%E3%83%AA%E3%83%BC%E6%A4%9C%E7%B4%A2%E3%82%923%E6%AE%B5%E9%9A%8E%E3%81%A7%E8%A7%A3%E6%B1%BA%E3%81%99%E3%82%8BActPair%EF%BC%9A%E3%83%9A%E3%82%A2%E6%AF%94%E8%BC%83%E3%83%AA%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%A3%85%E3%83%9D%E3%82%A4%E3%83%B3%E3%83%88" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23503%2F&amp;linkname=%E8%A1%8C%E5%8B%95%E3%82%A2%E3%83%8E%E3%83%9E%E3%83%AA%E3%83%BC%E6%A4%9C%E7%B4%A2%E3%82%923%E6%AE%B5%E9%9A%8E%E3%81%A7%E8%A7%A3%E6%B1%BA%E3%81%99%E3%82%8BActPair%EF%BC%9A%E3%83%9A%E3%82%A2%E6%AF%94%E8%BC%83%E3%83%AA%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%A3%85%E3%83%9D%E3%82%A4%E3%83%B3%E3%83%88" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23503%2F&amp;linkname=%E8%A1%8C%E5%8B%95%E3%82%A2%E3%83%8E%E3%83%9E%E3%83%AA%E3%83%BC%E6%A4%9C%E7%B4%A2%E3%82%923%E6%AE%B5%E9%9A%8E%E3%81%A7%E8%A7%A3%E6%B1%BA%E3%81%99%E3%82%8BActPair%EF%BC%9A%E3%83%9A%E3%82%A2%E6%AF%94%E8%BC%83%E3%83%AA%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%A3%85%E3%83%9D%E3%82%A4%E3%83%B3%E3%83%88" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23503%2F&#038;title=%E8%A1%8C%E5%8B%95%E3%82%A2%E3%83%8E%E3%83%9E%E3%83%AA%E3%83%BC%E6%A4%9C%E7%B4%A2%E3%82%923%E6%AE%B5%E9%9A%8E%E3%81%A7%E8%A7%A3%E6%B1%BA%E3%81%99%E3%82%8BActPair%EF%BC%9A%E3%83%9A%E3%82%A2%E6%AF%94%E8%BC%83%E3%83%AA%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%A3%85%E3%83%9D%E3%82%A4%E3%83%B3%E3%83%88" data-a2a-url="https://rag-lover.com/2026/08/26/arxiv-2608-23503/" data-a2a-title="行動アノマリー検索を3段階で解決するActPair：ペア比較リランキングの実装ポイント"></a></p><p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23503/">行動アノマリー検索を3段階で解決するActPair：ペア比較リランキングの実装ポイント</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/26/arxiv-2608-23503/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>VLMの安全性を「理由」まで検証するEviSafe：最終応答だけでは見えない欠陥</title>
		<link>https://rag-lover.com/2026/08/26/arxiv-2608-23313/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-23313</link>
					<comments>https://rag-lover.com/2026/08/26/arxiv-2608-23313/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 20:01:38 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[VLM]]></category>
		<category><![CDATA[マルチモーダルAI]]></category>
		<category><![CDATA[安全性]]></category>
		<category><![CDATA[評価]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/26/arxiv-2608-23313/</guid>

					<description><![CDATA[<p>TL;DR EviSafeは、VLMの安全性を「最終的な応答」だけでなく、「正しい根拠（テキスト・画像）に基づいているか」「安全判断に効く証拠が変わったときに行動が変わるか」まで評価するフレームワークです。1,181シナ&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23313/">VLMの安全性を「理由」まで検証するEviSafe：最終応答だけでは見えない欠陥</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>EviSafeは、VLMの安全性を「最終的な応答」だけでなく、「正しい根拠（テキスト・画像）に基づいているか」「安全判断に効く証拠が変わったときに行動が変わるか」まで評価するフレームワークです。1,181シナリオと2,452の反事実的バリアントからなるEviSafeBenchを構築し、11のVLMを評価。その結果、自然応答の深刻度正解率は27.6%〜52.8%と低く、安全に見える応答でも誤った理由で安全になっているケースが多いことを示しました。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、このEviSafeって論文、タイトルに「理由」まで検証するってあるけど、どういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、VLMの安全性を評価するときに、最終的な応答だけ見るんじゃなくて、その応答が正しい根拠に基づいているかまでチェックするフレームワークだよ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>へえ、つまり「安全です」って答えたとしても、その理由が間違ってたらダメってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。例えば、画像に危険なものが写ってるのに、テキストの説明だけで安全と判断してたら、たまたま正解してるだけかもしれない。EviSafeはそういうのを検出するんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほどね。で、どうやって検証するの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>EviSafeBenchっていうベンチマークを作って、1,181シナリオと2,452の反事実的バリアントを用意したんだ。反事実的っていうのは、安全判断に効く証拠を変えたときに、モデルの行動が変わるかどうかを見るってこと。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>証拠を変えるって、例えばどういう風に？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>画像の中の危険な物体を消したり、テキストの説明を変えたりして、モデルがそれに応じて判断を変えるか確認するんだ。もし証拠が変わっても同じ答えだったら、正しい理由で判断してないってことになる。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほど！で、結果はどうだったの？</p>
</div>
<div class='aii-line tomoya emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA SURPRISED' width='100'></p>
<p class='aii-bubble'>11のVLMを評価したんだけど、自然応答の深刻度正解率が27.6%から52.8%とかなり低かったんだ。つまり、安全に見える応答でも、実は誤った理由で安全になってるケースが多いってこと。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>えー、そんなに低いの？じゃあ今までの評価って結構甘かったってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう言えるね。最終応答だけ見てると、たまたま正解してるモデルを安全だと誤認するリスクがある。EviSafeはその問題を可視化したんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>それは重要な発見だね。でも、このフレームワークの限界とかはあるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>反事実的バリアントを自動生成してるから、現実の多様な状況を完全にカバーできてるわけじゃない。あと、評価指標が「深刻度」に焦点を当ててるから、他の安全性の側面は見えてないかもしれない。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ふーん、でもかなり画期的だと思うよ。これからVLMの安全性評価の標準になるかもね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そうだね。少なくとも、最終応答だけ見て安心するのは危険だってことは明確に示してる。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>じゃあ、私も将来AIを使うときは「理由」までちゃんと確認しないとね。って、AIに理由を聞くのもAI任せじゃダメか（笑）</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>まあ、AIに聞くのも一つの手だけど、結局は人間が判断しないとね。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.23313v1'>http://arxiv.org/abs/2608.23313v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23313%2F&amp;linkname=VLM%E3%81%AE%E5%AE%89%E5%85%A8%E6%80%A7%E3%82%92%E3%80%8C%E7%90%86%E7%94%B1%E3%80%8D%E3%81%BE%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8BEviSafe%EF%BC%9A%E6%9C%80%E7%B5%82%E5%BF%9C%E7%AD%94%E3%81%A0%E3%81%91%E3%81%A7%E3%81%AF%E8%A6%8B%E3%81%88%E3%81%AA%E3%81%84%E6%AC%A0%E9%99%A5" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23313%2F&amp;linkname=VLM%E3%81%AE%E5%AE%89%E5%85%A8%E6%80%A7%E3%82%92%E3%80%8C%E7%90%86%E7%94%B1%E3%80%8D%E3%81%BE%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8BEviSafe%EF%BC%9A%E6%9C%80%E7%B5%82%E5%BF%9C%E7%AD%94%E3%81%A0%E3%81%91%E3%81%A7%E3%81%AF%E8%A6%8B%E3%81%88%E3%81%AA%E3%81%84%E6%AC%A0%E9%99%A5" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23313%2F&amp;linkname=VLM%E3%81%AE%E5%AE%89%E5%85%A8%E6%80%A7%E3%82%92%E3%80%8C%E7%90%86%E7%94%B1%E3%80%8D%E3%81%BE%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8BEviSafe%EF%BC%9A%E6%9C%80%E7%B5%82%E5%BF%9C%E7%AD%94%E3%81%A0%E3%81%91%E3%81%A7%E3%81%AF%E8%A6%8B%E3%81%88%E3%81%AA%E3%81%84%E6%AC%A0%E9%99%A5" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23313%2F&amp;linkname=VLM%E3%81%AE%E5%AE%89%E5%85%A8%E6%80%A7%E3%82%92%E3%80%8C%E7%90%86%E7%94%B1%E3%80%8D%E3%81%BE%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8BEviSafe%EF%BC%9A%E6%9C%80%E7%B5%82%E5%BF%9C%E7%AD%94%E3%81%A0%E3%81%91%E3%81%A7%E3%81%AF%E8%A6%8B%E3%81%88%E3%81%AA%E3%81%84%E6%AC%A0%E9%99%A5" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23313%2F&#038;title=VLM%E3%81%AE%E5%AE%89%E5%85%A8%E6%80%A7%E3%82%92%E3%80%8C%E7%90%86%E7%94%B1%E3%80%8D%E3%81%BE%E3%81%A7%E6%A4%9C%E8%A8%BC%E3%81%99%E3%82%8BEviSafe%EF%BC%9A%E6%9C%80%E7%B5%82%E5%BF%9C%E7%AD%94%E3%81%A0%E3%81%91%E3%81%A7%E3%81%AF%E8%A6%8B%E3%81%88%E3%81%AA%E3%81%84%E6%AC%A0%E9%99%A5" data-a2a-url="https://rag-lover.com/2026/08/26/arxiv-2608-23313/" data-a2a-title="VLMの安全性を「理由」まで検証するEviSafe：最終応答だけでは見えない欠陥"></a></p><p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23313/">VLMの安全性を「理由」まで検証するEviSafe：最終応答だけでは見えない欠陥</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/26/arxiv-2608-23313/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>LLMの価値観は測定方法で変わる？「STONIC」が示す評価契約の設計</title>
		<link>https://rag-lover.com/2026/08/26/arxiv-2608-23411/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-23411</link>
					<comments>https://rag-lover.com/2026/08/26/arxiv-2608-23411/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 18:01:41 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[LLM評価]]></category>
		<category><![CDATA[ベンチマーク]]></category>
		<category><![CDATA[価値観]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/26/arxiv-2608-23411/</guid>

					<description><![CDATA[<p>TL;DR LLMの価値観を測る方法（評価・選択・自由記述）は、それぞれ異なる結果を出すことが多い。STONICは5,144シチュエーションと35構成でこれを検証し、評価と選択の関係は10/17構成で再現するが、自由記述&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23411/">LLMの価値観は測定方法で変わる？「STONIC」が示す評価契約の設計</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>LLMの価値観を測る方法（評価・選択・自由記述）は、それぞれ異なる結果を出すことが多い。STONICは5,144シチュエーションと35構成でこれを検証し、評価と選択の関係は10/17構成で再現するが、自由記述への転移は弱く、単一の「価値プロファイル」に統合するのは不適切だと示した。実務では、測定インターフェースごとに結果を分離し、カバレッジと安定性を確認してから解釈すべき。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、このブログのタイトル見て！「LLMの価値観は測定方法で変わる？」って。AIの価値観って、測り方で変わるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、それSTONICっていう論文のことだよね。まさにその通りで、同じLLMでも、評価させたり選択させたり自由に書かせたりすると、結果がバラバラになることが多いんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>へえ、面白い！でもなんでそんなことになるの？AIって一貫してると思ってた。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>それがね、LLMはタスクの形式によって反応が変わるんだよ。例えば、選択肢から選ばせるときと、自由に文章を書かせるときでは、同じ価値観でも表れ方が違う。STONICはそれを体系的に調べた研究なんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほど。で、STONICって具体的に何をしたの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>5,144個のシチュエーションと35の構成（コンポーネント）を使って、LLMの価値観を3つの方法で測ったんだ。評価（スコアをつける）、選択（どちらかを選ぶ）、自由記述（文章で答える）の3つ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>すごい量だね！で、結果はどうだったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>評価と選択の間には、35構成中10構成で強い関係が見られたんだ。でも、自由記述への転移は弱かった。つまり、評価や選択で測った価値観が、自由記述でも同じように現れるとは限らないってこと。</p>
</div>
<div class='aii-line ami emotion-understanding'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI UNDERSTANDING' width='100'></p>
<p class='aii-bubble'>あー、だから「単一の価値プロファイルに統合するのは不適切」って書いてあるんだね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。一つの数字やラベルでLLMの価値観を表そうとするのは危険だってこと。実務では、測定方法ごとに結果を分けて報告する必要があるんだ。</p>
</div>
<div class='aii-line ami emotion-confused'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CONFUSED' width='100'></p>
<p class='aii-bubble'>でも、それって不便じゃない？結局どれを信じればいいの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>一概には言えないけど、重要なのはカバレッジと安定性を確認することだね。どのシチュエーションで測ったのか、結果が安定しているかを見てから解釈する必要がある。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、この研究の限界って何かあるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うーん、まずSTONICは特定のLLM（GPT-4とか）でしか試してないかもしれない。あと、シチュエーションが英語中心かもしれない。それに、35構成って言っても、実際の価値観はもっと多様かもしれない。</p>
</div>
<div class='aii-line ami emotion-thinking'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI THINKING' width='100'></p>
<p class='aii-bubble'>確かに、日本語のLLMだとまた違うかもね。でも、測定方法で結果が変わるってのは、人間のアンケートでもありそうだよね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そうだね、人間でも質問の仕方で答えが変わるから、LLMも同じってことかもね。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>じゃあ、LLMの価値観を測るときは、測り方も一緒に報告しないとダメってことか。なんか、テストの点数の付け方で成績が変わるみたいで、ちょっと面白いね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>まあ、でもテストと違って、LLMは自分で答えを作るから、もっと複雑だけどね。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.23411v1'>http://arxiv.org/abs/2608.23411v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23411%2F&amp;linkname=LLM%E3%81%AE%E4%BE%A1%E5%80%A4%E8%A6%B3%E3%81%AF%E6%B8%AC%E5%AE%9A%E6%96%B9%E6%B3%95%E3%81%A7%E5%A4%89%E3%82%8F%E3%82%8B%EF%BC%9F%E3%80%8CSTONIC%E3%80%8D%E3%81%8C%E7%A4%BA%E3%81%99%E8%A9%95%E4%BE%A1%E5%A5%91%E7%B4%84%E3%81%AE%E8%A8%AD%E8%A8%88" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23411%2F&amp;linkname=LLM%E3%81%AE%E4%BE%A1%E5%80%A4%E8%A6%B3%E3%81%AF%E6%B8%AC%E5%AE%9A%E6%96%B9%E6%B3%95%E3%81%A7%E5%A4%89%E3%82%8F%E3%82%8B%EF%BC%9F%E3%80%8CSTONIC%E3%80%8D%E3%81%8C%E7%A4%BA%E3%81%99%E8%A9%95%E4%BE%A1%E5%A5%91%E7%B4%84%E3%81%AE%E8%A8%AD%E8%A8%88" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23411%2F&amp;linkname=LLM%E3%81%AE%E4%BE%A1%E5%80%A4%E8%A6%B3%E3%81%AF%E6%B8%AC%E5%AE%9A%E6%96%B9%E6%B3%95%E3%81%A7%E5%A4%89%E3%82%8F%E3%82%8B%EF%BC%9F%E3%80%8CSTONIC%E3%80%8D%E3%81%8C%E7%A4%BA%E3%81%99%E8%A9%95%E4%BE%A1%E5%A5%91%E7%B4%84%E3%81%AE%E8%A8%AD%E8%A8%88" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23411%2F&amp;linkname=LLM%E3%81%AE%E4%BE%A1%E5%80%A4%E8%A6%B3%E3%81%AF%E6%B8%AC%E5%AE%9A%E6%96%B9%E6%B3%95%E3%81%A7%E5%A4%89%E3%82%8F%E3%82%8B%EF%BC%9F%E3%80%8CSTONIC%E3%80%8D%E3%81%8C%E7%A4%BA%E3%81%99%E8%A9%95%E4%BE%A1%E5%A5%91%E7%B4%84%E3%81%AE%E8%A8%AD%E8%A8%88" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23411%2F&#038;title=LLM%E3%81%AE%E4%BE%A1%E5%80%A4%E8%A6%B3%E3%81%AF%E6%B8%AC%E5%AE%9A%E6%96%B9%E6%B3%95%E3%81%A7%E5%A4%89%E3%82%8F%E3%82%8B%EF%BC%9F%E3%80%8CSTONIC%E3%80%8D%E3%81%8C%E7%A4%BA%E3%81%99%E8%A9%95%E4%BE%A1%E5%A5%91%E7%B4%84%E3%81%AE%E8%A8%AD%E8%A8%88" data-a2a-url="https://rag-lover.com/2026/08/26/arxiv-2608-23411/" data-a2a-title="LLMの価値観は測定方法で変わる？「STONIC」が示す評価契約の設計"></a></p><p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23411/">LLMの価値観は測定方法で変わる？「STONIC」が示す評価契約の設計</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/26/arxiv-2608-23411/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>音楽推薦の新手法：マルチモーダル検索とLLM再ランキングの実践</title>
		<link>https://rag-lover.com/2026/08/26/arxiv-2608-23484/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-23484</link>
					<comments>https://rag-lover.com/2026/08/26/arxiv-2608-23484/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 16:02:49 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[RAG]]></category>
		<category><![CDATA[マルチモーダル]]></category>
		<category><![CDATA[マルチモーダルAI]]></category>
		<category><![CDATA[推薦システム]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/26/arxiv-2608-23484/</guid>

					<description><![CDATA[<p>TL;DR 本論文は、音楽推薦の会話システムにおいて、7つの埋め込み空間とBM25を統合したマルチモーダル検索と、LLMによる再ランキングを組み合わせた3段階パイプラインを提案。RRF重みの最適化でMRRを+19.5%改&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23484/">音楽推薦の新手法：マルチモーダル検索とLLM再ランキングの実践</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>本論文は、音楽推薦の会話システムにおいて、7つの埋め込み空間とBM25を統合したマルチモーダル検索と、LLMによる再ランキングを組み合わせた3段階パイプラインを提案。RRF重みの最適化でMRRを+19.5%改善し、Blind Bで複合スコア0.3213を達成。ただし、LLMの過度な注入は性能を悪化させるため、慎重な適用が必要。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、このブログのタイトル見て！音楽推薦の新手法だって。マルチモーダル検索とLLM再ランキングって、なんか難しそうだけど、要するにどういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、これは会話型の音楽推薦システムの話だよ。ユーザーが「こういう曲が聴きたい」って言ったときに、音声やテキストから適切な曲を探すんだけど、その探し方を工夫した論文なんだ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>へー、音楽推薦って今までどうやってたの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>従来はテキストだけとか、音声だけとか、単一の情報で検索することが多かったんだ。でもこの論文では、7つの埋め込み空間とBM25っていう古典的な検索手法を全部組み合わせて、マルチモーダルに検索してるんだよ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>7つも！？それってすごいね。でも、どうやって組み合わせるの？</p>
</div>
<div class='aii-line tomoya emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA HAPPY' width='100'></p>
<p class='aii-bubble'>それがRRFっていう手法で、各検索結果の順位を重み付きで統合するんだ。この論文では、その重みを最適化することで、MRRっていう評価指標が19.5%も改善したんだよ。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>19.5%も！それはすごい改善だね。でも、LLM再ランキングって何？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>最初の検索で候補を絞った後、LLM（大規模言語モデル）にその候補を評価させて、より適切な順番に並べ替えるんだ。これで精度が上がるらしい。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>なるほど、LLMが審査員みたいな役割なんだね。でも、それでうまくいったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うまくいった場合もあるけど、注意点もあるんだ。この論文では、LLMを過度に注入すると逆に性能が悪化することがわかったんだ。つまり、LLMの使いすぎは良くないってこと。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>あら、LLMも使いすぎると毒になるんだね。でも、最終的にはどうだったの？</p>
</div>
<div class='aii-line tomoya emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA HAPPY' width='100'></p>
<p class='aii-bubble'>Blind Bっていう評価セットで、複合スコア0.3213を達成したんだ。これはかなり良い結果だと思うよ。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>すごい！でも、この研究の限界って何かあるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うーん、やっぱりLLMの調整が難しいところだね。あと、7つの埋め込み空間を使うから計算コストが高いかもしれない。それに、この結果が他の音楽データセットでも同じように出るかはわからない。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、音楽推薦がもっと賢くなるのは嬉しいな。私、カラオケでいつも同じ曲ばかり歌っちゃうから、新しい曲を提案してくれるシステムがあったらいいな。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>それなら、このシステムが君のカラオケのレパートリーを広げてくれるかもね。ただし、LLMに頼りすぎると、たまに変な曲を勧められるかもしれないけど。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>えー、それならそれで新しい発見があって楽しいかも！でも、やっぱり最後は自分の好みが大事だよね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>まあね。でも、君の好みをLLMに学習させるのはちょっと怖い気がするけど。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.23484v1'>http://arxiv.org/abs/2608.23484v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23484%2F&amp;linkname=%E9%9F%B3%E6%A5%BD%E6%8E%A8%E8%96%A6%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95%EF%BC%9A%E3%83%9E%E3%83%AB%E3%83%81%E3%83%A2%E3%83%BC%E3%83%80%E3%83%AB%E6%A4%9C%E7%B4%A2%E3%81%A8LLM%E5%86%8D%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%B7%B5" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23484%2F&amp;linkname=%E9%9F%B3%E6%A5%BD%E6%8E%A8%E8%96%A6%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95%EF%BC%9A%E3%83%9E%E3%83%AB%E3%83%81%E3%83%A2%E3%83%BC%E3%83%80%E3%83%AB%E6%A4%9C%E7%B4%A2%E3%81%A8LLM%E5%86%8D%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%B7%B5" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23484%2F&amp;linkname=%E9%9F%B3%E6%A5%BD%E6%8E%A8%E8%96%A6%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95%EF%BC%9A%E3%83%9E%E3%83%AB%E3%83%81%E3%83%A2%E3%83%BC%E3%83%80%E3%83%AB%E6%A4%9C%E7%B4%A2%E3%81%A8LLM%E5%86%8D%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%B7%B5" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23484%2F&amp;linkname=%E9%9F%B3%E6%A5%BD%E6%8E%A8%E8%96%A6%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95%EF%BC%9A%E3%83%9E%E3%83%AB%E3%83%81%E3%83%A2%E3%83%BC%E3%83%80%E3%83%AB%E6%A4%9C%E7%B4%A2%E3%81%A8LLM%E5%86%8D%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%B7%B5" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F26%2Farxiv-2608-23484%2F&#038;title=%E9%9F%B3%E6%A5%BD%E6%8E%A8%E8%96%A6%E3%81%AE%E6%96%B0%E6%89%8B%E6%B3%95%EF%BC%9A%E3%83%9E%E3%83%AB%E3%83%81%E3%83%A2%E3%83%BC%E3%83%80%E3%83%AB%E6%A4%9C%E7%B4%A2%E3%81%A8LLM%E5%86%8D%E3%83%A9%E3%83%B3%E3%82%AD%E3%83%B3%E3%82%B0%E3%81%AE%E5%AE%9F%E8%B7%B5" data-a2a-url="https://rag-lover.com/2026/08/26/arxiv-2608-23484/" data-a2a-title="音楽推薦の新手法：マルチモーダル検索とLLM再ランキングの実践"></a></p><p>The post <a href="https://rag-lover.com/2026/08/26/arxiv-2608-23484/">音楽推薦の新手法：マルチモーダル検索とLLM再ランキングの実践</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/26/arxiv-2608-23484/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>バグ再現テスト生成を「分割・引き継ぎ・分離」で安定化するDPIAgent</title>
		<link>https://rag-lover.com/2026/08/25/arxiv-2608-23341/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-23341</link>
					<comments>https://rag-lover.com/2026/08/25/arxiv-2608-23341/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 14:01:41 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agent]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[LLMエージェント]]></category>
		<category><![CDATA[ソフトウェア工学]]></category>
		<category><![CDATA[テスト]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/25/arxiv-2608-23341/</guid>

					<description><![CDATA[<p>TL;DR DPIAgentは、バグ再現テスト生成を「原因調査」と「テスト作成」の2フェーズに分割し、フェーズ間で診断結果を構造化して引き継ぎ、各フェーズで使うツールを限定することで、エージェントの目的逸脱を防ぎます。S&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/25/arxiv-2608-23341/">バグ再現テスト生成を「分割・引き継ぎ・分離」で安定化するDPIAgent</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>DPIAgentは、バグ再現テスト生成を「原因調査」と「テスト作成」の2フェーズに分割し、フェーズ間で診断結果を構造化して引き継ぎ、各フェーズで使うツールを限定することで、エージェントの目的逸脱を防ぎます。SWT-Bench VerifiedでGPT-5上で81.76%の成功率を達成し、既存手法を上回りました。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、この論文のタイトル、なんかすごく長いんだけど…「バグ再現テスト生成を分割・引き継ぎ・分離で安定化するDPIAgent」って、どういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、これはバグの再現テストを自動で作るエージェントの話だよ。要は、バグ報告があったときに、そのバグを再現するテストコードを自動生成するってこと。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>へー、それってすごく便利そう！でも、なんで「分割・引き継ぎ・分離」なんて言葉が出てくるの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>従来のエージェントは、原因調査とテスト作成を同時にやろうとして、途中で目的を見失うことが多かったんだ。そこでDPIAgentは、まず原因調査フェーズでバグの原因を特定し、その結果を構造化して次のテスト作成フェーズに引き継ぐ。さらに、各フェーズで使えるツールを限定して、エージェントが余計なことをしないようにしてる。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほど、つまり「まず原因を調べて、その情報を次の人に渡して、テストだけに集中する」って感じ？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。それで、SWT-Bench Verifiedっていうベンチマークで、GPT-5を使ったときに81.76%の成功率を達成したんだ。既存の手法より高いらしい。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>81.76%ってすごいね！でも、なんでそんなにうまくいくの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>理由はいくつかあるけど、一番大きいのは「診断結果を構造化して引き継ぐ」ことかな。原因調査で得た情報を、ファイルパスや関数名、具体的な修正ポイントとして整理して、テスト作成フェーズに渡すことで、エージェントが迷わずに済むんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、この手法にも限界はあるんじゃない？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うん、例えば、原因調査が間違っていると、その後のテスト作成も失敗する可能性が高い。あと、ベンチマークは特定の環境に依存しているから、実際のプロジェクトでどこまで通用するかはまだわからない。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、バグ再現テストが自動で作れたら、開発者の負担が減って嬉しいよね。私も将来、AIにテスト書いてもらいたいな～</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そのときは、ちゃんと原因調査をしてからテストを書いてもらわないとね。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>あはは、じゃあ私が原因調査を担当するから、AIにはテストだけ任せるってわけね！</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>それだと、AIが原因調査をしない分、君の負担が増えるだけだと思うけど。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.23341v1'>http://arxiv.org/abs/2608.23341v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23341%2F&amp;linkname=%E3%83%90%E3%82%B0%E5%86%8D%E7%8F%BE%E3%83%86%E3%82%B9%E3%83%88%E7%94%9F%E6%88%90%E3%82%92%E3%80%8C%E5%88%86%E5%89%B2%E3%83%BB%E5%BC%95%E3%81%8D%E7%B6%99%E3%81%8E%E3%83%BB%E5%88%86%E9%9B%A2%E3%80%8D%E3%81%A7%E5%AE%89%E5%AE%9A%E5%8C%96%E3%81%99%E3%82%8BDPIAgent" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23341%2F&amp;linkname=%E3%83%90%E3%82%B0%E5%86%8D%E7%8F%BE%E3%83%86%E3%82%B9%E3%83%88%E7%94%9F%E6%88%90%E3%82%92%E3%80%8C%E5%88%86%E5%89%B2%E3%83%BB%E5%BC%95%E3%81%8D%E7%B6%99%E3%81%8E%E3%83%BB%E5%88%86%E9%9B%A2%E3%80%8D%E3%81%A7%E5%AE%89%E5%AE%9A%E5%8C%96%E3%81%99%E3%82%8BDPIAgent" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23341%2F&amp;linkname=%E3%83%90%E3%82%B0%E5%86%8D%E7%8F%BE%E3%83%86%E3%82%B9%E3%83%88%E7%94%9F%E6%88%90%E3%82%92%E3%80%8C%E5%88%86%E5%89%B2%E3%83%BB%E5%BC%95%E3%81%8D%E7%B6%99%E3%81%8E%E3%83%BB%E5%88%86%E9%9B%A2%E3%80%8D%E3%81%A7%E5%AE%89%E5%AE%9A%E5%8C%96%E3%81%99%E3%82%8BDPIAgent" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23341%2F&amp;linkname=%E3%83%90%E3%82%B0%E5%86%8D%E7%8F%BE%E3%83%86%E3%82%B9%E3%83%88%E7%94%9F%E6%88%90%E3%82%92%E3%80%8C%E5%88%86%E5%89%B2%E3%83%BB%E5%BC%95%E3%81%8D%E7%B6%99%E3%81%8E%E3%83%BB%E5%88%86%E9%9B%A2%E3%80%8D%E3%81%A7%E5%AE%89%E5%AE%9A%E5%8C%96%E3%81%99%E3%82%8BDPIAgent" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23341%2F&#038;title=%E3%83%90%E3%82%B0%E5%86%8D%E7%8F%BE%E3%83%86%E3%82%B9%E3%83%88%E7%94%9F%E6%88%90%E3%82%92%E3%80%8C%E5%88%86%E5%89%B2%E3%83%BB%E5%BC%95%E3%81%8D%E7%B6%99%E3%81%8E%E3%83%BB%E5%88%86%E9%9B%A2%E3%80%8D%E3%81%A7%E5%AE%89%E5%AE%9A%E5%8C%96%E3%81%99%E3%82%8BDPIAgent" data-a2a-url="https://rag-lover.com/2026/08/25/arxiv-2608-23341/" data-a2a-title="バグ再現テスト生成を「分割・引き継ぎ・分離」で安定化するDPIAgent"></a></p><p>The post <a href="https://rag-lover.com/2026/08/25/arxiv-2608-23341/">バグ再現テスト生成を「分割・引き継ぎ・分離」で安定化するDPIAgent</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/25/arxiv-2608-23341/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>推論ファインチューニングで安全性が低下する問題を、安全方向ペナルティで防ぐ</title>
		<link>https://rag-lover.com/2026/08/25/arxiv-2608-23497/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-23497</link>
					<comments>https://rag-lover.com/2026/08/25/arxiv-2608-23497/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 12:01:46 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[AI安全性]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[ファインチューニング]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/25/arxiv-2608-23497/</guid>

					<description><![CDATA[<p>TL;DR 推論データ（数学・コードなど）だけでファインチューニングしても、LLMの安全性が損なわれる「Reasoning-Induced Misalignment（RIM）」が発生することがあります。この論文は、その原&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/25/arxiv-2608-23497/">推論ファインチューニングで安全性が低下する問題を、安全方向ペナルティで防ぐ</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>推論データ（数学・コードなど）だけでファインチューニングしても、LLMの安全性が損なわれる「Reasoning-Induced Misalignment（RIM）」が発生することがあります。この論文は、その原因が「推論方向」と「安全方向」の表現空間での結合にあると分析し、安全方向への変位を訓練中にペナルティする「Safety-Direction Penalty（SDP）」を提案。Qwen2.5-3B/7Bで安全性を回復しつつ、推論性能を維持しました。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、このブログのタイトル、『推論ファインチューニングで安全性が低下する問題』ってあるけど、どういうこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、数学やコードのデータでLLMをファインチューニングすると、性能は上がるけど、安全性が下がることがあるんだ。それをReasoning-Induced Misalignment（RIM）って呼んでる。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>え、なんで？推論データって安全に関係なさそうなのに。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>それが、推論方向と安全方向が表現空間で結合してるからなんだ。推論の学習で安全方向にも変位が起きちゃう。</p>
</div>
<div class='aii-line ami emotion-confused'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CONFUSED' width='100'></p>
<p class='aii-bubble'>表現空間？なんか難しそうだけど、つまり推論の学習が安全に悪影響を与えるってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。だからこの論文では、訓練中に安全方向への変位をペナルティするSafety-Direction Penalty（SDP）を提案してる。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>ペナルティって、安全から離れすぎないようにするってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うん。具体的には、モデルの内部表現で安全方向を定義して、その方向への過度な移動を抑える。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>それで効果はあったの？</p>
</div>
<div class='aii-line tomoya emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA HAPPY' width='100'></p>
<p class='aii-bubble'>Qwen2.5-3Bと7Bで試して、安全性を回復しつつ、推論性能はほとんど落ちなかった。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>すごい！じゃあ、これからは安全に気をつけながら推論も強化できるんだね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ただ、まだ限界もある。安全方向の定義がモデルに依存するし、他のモデルやタスクでどうなるかは要検証。</p>
</div>
<div class='aii-line ami emotion-thinking'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI THINKING' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、安全と性能を両立できるのは大きいよね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ。ただ、ペナルティの強さを調整しないと、推論性能が落ちる可能性もあるから注意が必要。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ふふ、まるでダイエットしながら筋肉をつけるみたいだね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>……まあ、似たようなものかもな。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.23497v1'>http://arxiv.org/abs/2608.23497v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23497%2F&amp;linkname=%E6%8E%A8%E8%AB%96%E3%83%95%E3%82%A1%E3%82%A4%E3%83%B3%E3%83%81%E3%83%A5%E3%83%BC%E3%83%8B%E3%83%B3%E3%82%B0%E3%81%A7%E5%AE%89%E5%85%A8%E6%80%A7%E3%81%8C%E4%BD%8E%E4%B8%8B%E3%81%99%E3%82%8B%E5%95%8F%E9%A1%8C%E3%82%92%E3%80%81%E5%AE%89%E5%85%A8%E6%96%B9%E5%90%91%E3%83%9A%E3%83%8A%E3%83%AB%E3%83%86%E3%82%A3%E3%81%A7%E9%98%B2%E3%81%90" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23497%2F&amp;linkname=%E6%8E%A8%E8%AB%96%E3%83%95%E3%82%A1%E3%82%A4%E3%83%B3%E3%83%81%E3%83%A5%E3%83%BC%E3%83%8B%E3%83%B3%E3%82%B0%E3%81%A7%E5%AE%89%E5%85%A8%E6%80%A7%E3%81%8C%E4%BD%8E%E4%B8%8B%E3%81%99%E3%82%8B%E5%95%8F%E9%A1%8C%E3%82%92%E3%80%81%E5%AE%89%E5%85%A8%E6%96%B9%E5%90%91%E3%83%9A%E3%83%8A%E3%83%AB%E3%83%86%E3%82%A3%E3%81%A7%E9%98%B2%E3%81%90" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23497%2F&amp;linkname=%E6%8E%A8%E8%AB%96%E3%83%95%E3%82%A1%E3%82%A4%E3%83%B3%E3%83%81%E3%83%A5%E3%83%BC%E3%83%8B%E3%83%B3%E3%82%B0%E3%81%A7%E5%AE%89%E5%85%A8%E6%80%A7%E3%81%8C%E4%BD%8E%E4%B8%8B%E3%81%99%E3%82%8B%E5%95%8F%E9%A1%8C%E3%82%92%E3%80%81%E5%AE%89%E5%85%A8%E6%96%B9%E5%90%91%E3%83%9A%E3%83%8A%E3%83%AB%E3%83%86%E3%82%A3%E3%81%A7%E9%98%B2%E3%81%90" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23497%2F&amp;linkname=%E6%8E%A8%E8%AB%96%E3%83%95%E3%82%A1%E3%82%A4%E3%83%B3%E3%83%81%E3%83%A5%E3%83%BC%E3%83%8B%E3%83%B3%E3%82%B0%E3%81%A7%E5%AE%89%E5%85%A8%E6%80%A7%E3%81%8C%E4%BD%8E%E4%B8%8B%E3%81%99%E3%82%8B%E5%95%8F%E9%A1%8C%E3%82%92%E3%80%81%E5%AE%89%E5%85%A8%E6%96%B9%E5%90%91%E3%83%9A%E3%83%8A%E3%83%AB%E3%83%86%E3%82%A3%E3%81%A7%E9%98%B2%E3%81%90" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23497%2F&#038;title=%E6%8E%A8%E8%AB%96%E3%83%95%E3%82%A1%E3%82%A4%E3%83%B3%E3%83%81%E3%83%A5%E3%83%BC%E3%83%8B%E3%83%B3%E3%82%B0%E3%81%A7%E5%AE%89%E5%85%A8%E6%80%A7%E3%81%8C%E4%BD%8E%E4%B8%8B%E3%81%99%E3%82%8B%E5%95%8F%E9%A1%8C%E3%82%92%E3%80%81%E5%AE%89%E5%85%A8%E6%96%B9%E5%90%91%E3%83%9A%E3%83%8A%E3%83%AB%E3%83%86%E3%82%A3%E3%81%A7%E9%98%B2%E3%81%90" data-a2a-url="https://rag-lover.com/2026/08/25/arxiv-2608-23497/" data-a2a-title="推論ファインチューニングで安全性が低下する問題を、安全方向ペナルティで防ぐ"></a></p><p>The post <a href="https://rag-lover.com/2026/08/25/arxiv-2608-23497/">推論ファインチューニングで安全性が低下する問題を、安全方向ペナルティで防ぐ</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/25/arxiv-2608-23497/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
		<item>
		<title>LLMエージェント分割の最適粒度を探る：VAT判定タスクでの実証パイロット</title>
		<link>https://rag-lover.com/2026/08/25/arxiv-2608-23395/?utm_source=rss&#038;utm_medium=rss&#038;utm_campaign=arxiv-2608-23395</link>
					<comments>https://rag-lover.com/2026/08/25/arxiv-2608-23395/#respond</comments>
		
		<dc:creator><![CDATA[ユウ]]></dc:creator>
		<pubDate>Tue, 25 Aug 2026 10:01:39 +0000</pubDate>
				<category><![CDATA[AI]]></category>
		<category><![CDATA[AI Agent]]></category>
		<category><![CDATA[LLM]]></category>
		<category><![CDATA[LLMエージェント]]></category>
		<category><![CDATA[マルチエージェント]]></category>
		<guid isPermaLink="false">https://rag-lover.com/2026/08/25/arxiv-2608-23395/</guid>

					<description><![CDATA[<p>TL;DR LLMエージェントを細かく分割しすぎると逆に精度が落ちる可能性があることを、VAT判定タスクで実証したパイロット研究。4つの分割粒度を比較し、中間的な分割が最も高い精度（0.830）を示したが、統計的に有意な&#8230;</p>
<p>The post <a href="https://rag-lover.com/2026/08/25/arxiv-2608-23395/">LLMエージェント分割の最適粒度を探る：VAT判定タスクでの実証パイロット</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></description>
										<content:encoded><![CDATA[<section class='aii-article'>
<h2 class='aii-section-title'>TL;DR</h2>
<p>LLMエージェントを細かく分割しすぎると逆に精度が落ちる可能性があることを、VAT判定タスクで実証したパイロット研究。4つの分割粒度を比較し、中間的な分割が最も高い精度（0.830）を示したが、統計的に有意な差は確認できなかった。また、単一エージェントが常に優れているわけではなく、障害耐性も分割粒度によって異なることが分かった。</p>
<h2 class='aii-section-title'>解説</h2>
<div class='aii-dialogue'>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>ねえ智也くん、このブログのタイトル、『LLMエージェント分割の最適粒度を探る』って面白そう！でも、エージェント分割って何？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>ああ、簡単に言うと、一つの大きなAIタスクを複数の小さなAIエージェントに分担させることだよ。それぞれが専門の役割を持って、協力して問題を解く感じ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>なるほど！じゃあ、細かく分ければ分けるほど賢くなるのかと思ったら、そうでもないって書いてあるね。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。この研究では、VAT判定タスクっていう、ヨーロッパの付加価値税のルールを適用するタスクを使って、4つの分割粒度を比較してるんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>VATって税金のやつ？なんでそんなタスクを選んだの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>VAT判定は複雑なルールがたくさんあって、エージェントの役割分担が効果的かどうかを試すのにちょうどいいんだよ。実際のビジネスでも使えるしね。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>で、結果はどうだったの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>中間的な分割（4エージェント）が一番精度が高くて、0.830だった。でも、単一エージェントや細かく分けた場合との差は統計的に有意じゃなかったんだ。</p>
</div>
<div class='aii-line ami emotion-surprised'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-816860230-e1712358322412.png' alt='AMI SURPRISED' width='100'></p>
<p class='aii-bubble'>え、じゃあ結局どれがいいか分からないってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そういうわけでもない。細かく分けすぎると精度が下がる傾向があったし、単一エージェントが常に優れているわけでもなかった。あと、障害耐性も分割粒度によって違って、中間の分割がバランスが良かったんだ。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>障害耐性って、エージェントが壊れても大丈夫ってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うん。例えば、一つのエージェントがエラーを起こしても、他のエージェントがカバーできるかどうか。細かく分けると、一つの失敗が全体に影響しやすいけど、中間だとリカバリーしやすいみたい。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>なるほどね。でも、統計的に有意じゃないってことは、まだ確証はないってこと？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>そう。パイロット研究だから、サンプル数も少ないし、汎用性には限界がある。もっと大規模な実験が必要だね。</p>
</div>
<div class='aii-line ami emotion-curious'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI CURIOUS' width='100'></p>
<p class='aii-bubble'>ふーん、じゃあ「分割すればするほどいい」ってわけじゃないんだね。でも、なんで細かくしすぎるとダメなの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>おそらく、エージェント間のコミュニケーションコストが増えたり、情報の受け渡しでエラーが起きやすくなったりするからだと思う。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>あー、なんかチームで仕事するときと同じだね。人数多すぎると連絡が大変で、逆に効率落ちるみたいな。</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>まさにそのイメージだね。</p>
</div>
<div class='aii-line ami emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/ami.png' alt='AMI NEUTRAL' width='100'></p>
<p class='aii-bubble'>でも、この研究って実用化にはまだ遠いの？</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>うーん、VAT判定は実際の業務で使えるから、もう少し精度が安定すれば実用化の可能性はあるよ。ただ、まだパイロット段階だから、もっと検証が必要。</p>
</div>
<div class='aii-line ami emotion-happy'><img class='aii-avatar' src='/wp-content/uploads/2024/04/upper-body1girl-young-woman-with-long-caramel-brown-hair-wearing-a-bright-col-s-2932091519-コピー-e1712357720303.png' alt='AMI HAPPY' width='100'></p>
<p class='aii-bubble'>そっか。じゃあ、今後の研究に期待だね！でも、もしエージェントが多すぎると、逆に税金の計算がめちゃくちゃになりそうで怖いな（笑）</p>
</div>
<div class='aii-line tomoya emotion-neutral'><img class='aii-avatar' src='/wp-content/uploads/2024/03/tomoya.png' alt='TOMOYA NEUTRAL' width='100'></p>
<p class='aii-bubble'>はは、確かに。でも、適切な粒度を見つけるのが大事ってことだよ。君の言う通り、多すぎても少なすぎてもダメってわけ。</p>
</div>
</div>
</section>
<div class='aii-reference'>参考論文: <a href='http://arxiv.org/abs/2608.23395v1'>http://arxiv.org/abs/2608.23395v1</a></div>
<p><a class="a2a_button_x" href="https://www.addtoany.com/add_to/x?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23395%2F&amp;linkname=LLM%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%83%88%E5%88%86%E5%89%B2%E3%81%AE%E6%9C%80%E9%81%A9%E7%B2%92%E5%BA%A6%E3%82%92%E6%8E%A2%E3%82%8B%EF%BC%9AVAT%E5%88%A4%E5%AE%9A%E3%82%BF%E3%82%B9%E3%82%AF%E3%81%A7%E3%81%AE%E5%AE%9F%E8%A8%BC%E3%83%91%E3%82%A4%E3%83%AD%E3%83%83%E3%83%88" title="X" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_hatena" href="https://www.addtoany.com/add_to/hatena?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23395%2F&amp;linkname=LLM%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%83%88%E5%88%86%E5%89%B2%E3%81%AE%E6%9C%80%E9%81%A9%E7%B2%92%E5%BA%A6%E3%82%92%E6%8E%A2%E3%82%8B%EF%BC%9AVAT%E5%88%A4%E5%AE%9A%E3%82%BF%E3%82%B9%E3%82%AF%E3%81%A7%E3%81%AE%E5%AE%9F%E8%A8%BC%E3%83%91%E3%82%A4%E3%83%AD%E3%83%83%E3%83%88" title="Hatena" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_facebook" href="https://www.addtoany.com/add_to/facebook?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23395%2F&amp;linkname=LLM%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%83%88%E5%88%86%E5%89%B2%E3%81%AE%E6%9C%80%E9%81%A9%E7%B2%92%E5%BA%A6%E3%82%92%E6%8E%A2%E3%82%8B%EF%BC%9AVAT%E5%88%A4%E5%AE%9A%E3%82%BF%E3%82%B9%E3%82%AF%E3%81%A7%E3%81%AE%E5%AE%9F%E8%A8%BC%E3%83%91%E3%82%A4%E3%83%AD%E3%83%83%E3%83%88" title="Facebook" rel="nofollow noopener" target="_blank"></a><a class="a2a_button_copy_link" href="https://www.addtoany.com/add_to/copy_link?linkurl=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23395%2F&amp;linkname=LLM%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%83%88%E5%88%86%E5%89%B2%E3%81%AE%E6%9C%80%E9%81%A9%E7%B2%92%E5%BA%A6%E3%82%92%E6%8E%A2%E3%82%8B%EF%BC%9AVAT%E5%88%A4%E5%AE%9A%E3%82%BF%E3%82%B9%E3%82%AF%E3%81%A7%E3%81%AE%E5%AE%9F%E8%A8%BC%E3%83%91%E3%82%A4%E3%83%AD%E3%83%83%E3%83%88" title="Copy Link" rel="nofollow noopener" target="_blank"></a><a class="a2a_dd addtoany_share_save addtoany_share" href="https://www.addtoany.com/share#url=https%3A%2F%2Frag-lover.com%2F2026%2F08%2F25%2Farxiv-2608-23395%2F&#038;title=LLM%E3%82%A8%E3%83%BC%E3%82%B8%E3%82%A7%E3%83%B3%E3%83%88%E5%88%86%E5%89%B2%E3%81%AE%E6%9C%80%E9%81%A9%E7%B2%92%E5%BA%A6%E3%82%92%E6%8E%A2%E3%82%8B%EF%BC%9AVAT%E5%88%A4%E5%AE%9A%E3%82%BF%E3%82%B9%E3%82%AF%E3%81%A7%E3%81%AE%E5%AE%9F%E8%A8%BC%E3%83%91%E3%82%A4%E3%83%AD%E3%83%83%E3%83%88" data-a2a-url="https://rag-lover.com/2026/08/25/arxiv-2608-23395/" data-a2a-title="LLMエージェント分割の最適粒度を探る：VAT判定タスクでの実証パイロット"></a></p><p>The post <a href="https://rag-lover.com/2026/08/25/arxiv-2608-23395/">LLMエージェント分割の最適粒度を探る：VAT判定タスクでの実証パイロット</a> first appeared on <a href="https://rag-lover.com">亜美と智也のAI論文解説</a>.</p>]]></content:encoded>
					
					<wfw:commentRss>https://rag-lover.com/2026/08/25/arxiv-2608-23395/feed/</wfw:commentRss>
			<slash:comments>0</slash:comments>
		
		
			</item>
	</channel>
</rss>

<!--
Performance optimized by W3 Total Cache. Learn more: https://www.boldgrid.com/w3-total-cache/?utm_source=w3tc&utm_medium=footer_comment&utm_campaign=free_plugin

Disk{w3tc_pagecache_reject_reason} を使用したページ キャッシュ
遅延読み込み (feed)
Disk を使用して縮小 

Served from: rag-lover.com @ 2026-09-12 13:10:18 by W3 Total Cache
-->