Urgent.News

What's breaking now, across thousands of outlets.

AI

An LLM observability platform stores prompts, and prompts are the application

An LLM observability platform stores prompts, and prompts are the application A title query for Langfuse returns 346 matches in ZoomEye. The number is small and the contents are unusual. An observability tool for language models records the text that goes into them and the text that comes back, which makes a tracing store closer to a source repository than to a metrics backend. Context and method…

An LLM observability platform records the text inputs and outputs of language models, treating prompts as the application logic. Each language model call generates a trace containing the prompt, model parameters, completion, token counts, latency, and any attached metadata. Traces are sensitive as they hold whatever the application sends, potentially including secrets.

Langfuse indexes these traces, which can include production system prompts, outputs, and embedded secrets. To protect your organization, review all tracing infrastructure as a system of record for application text.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at dev.to →

More in AI

Tests are the only fixed point left when the implementation is disposable

Originally published at https://aicoding-guide.com . The more disposable the implementation becomes, the more the spec moves into the tests. That is the claim.

  • Tests serve as the fixed point in disposable implementations.
  • Tests guide correctness when code is rewritten repeatedly.
  • Certain domains require alternative fixed points beyond tests.

Your LLM Types One Token at a Time. It Doesn't Have To.

Every token your LLM emits costs one full forward pass through the entire model. Seventy billion parameters loaded from memory, multiplied, discarded — for a single token. Then again. And again.

  • Speculative decoding drafts tokens with cheap model before big model verification
  • Acceptance rule accepts drafted token with probability min(1, q(d)/p(d))
  • EAGLE-3 achieves 2-3x speedup over vanilla decoding, but diminishing returns at high batch sizes

The agent finished. Who turns off the VM?

Suppose a coding agent opens a pull request at 6 p.m. The tests pass, a preview is running, and the reviewer has already logged off. The agent's task is complete.

  • Agent session ends when task result is saved
  • Preview retained only until explicit expiry time
  • Workspace removal depends on commit status

More from Thursday 8 October →