Linked tool modules surround an illuminated task card on an engineering bench with an empty review chair behind it.
Conceptual artwork. Diagrams and examples below explain the technical details.

What changes when an AI system can act

A language model can explain an algorithm or draft a script. An agentic system can also request a tool operation, inspect the result and choose another operation. That change is useful in research because many tasks involve feedback: a file has an unexpected shape, a dependency fails, or a calculation reveals that a proposed operation is too large for memory. The system needs to respond to what actually happened.

For hyperspectral imaging, or HSI, a sensible starting point is a public-data assistant that checks metadata and prepares a small, reviewable analysis. The aim is a traceable research aid. It should state what it inspected, what it calculated and what remains uncertain. A fluent explanation is helpful, but the data, tool outputs and documented assumptions carry the scientific evidence.

A bounded assistant proposes actions; application controls determine execution.Task and budget to Proposed tool call: scope; Proposed tool call to Permission and parameter checks: request; Permission and parameter checks to Documented tool: allowed; Permission and parameter checks to Reject or stop: invalid or over budget; Documented tool to Result and run record: observation; Result and run record to Proposed tool call: feedbackConceptual relationshipsTask and budgetProposed tool callPermission andparameter checksDocumented toolResult and runrecordReject or stopscoperequestallowedinvalid or overbudgetobservationfeedback
Conceptual illustration. A bounded assistant proposes actions; application controls determine execution. Dashed side arrows show feedback, not an additional forward stage.

A model, a workflow and an agent are different components

A model supplies predictions or generated content. A workflow determines the sequence in which models and tools run. An agent allows the model to choose some of that sequence in response to observations. Anthropic’s published architectural distinction separates predefined workflows from systems where the model directs its process and tool use. This is a useful vocabulary, rather than a universal definition of intelligence.

ReAct is an early research example of combining language-model reasoning with actions that obtain information from an environment. Its reported experiments cover specific question-answering and interactive tasks. Those results establish a research precedent for feedback through tools; they do not establish that a language model can independently validate an HSI experiment. A researcher still needs suitable measurements, checks and domain judgement.

Sources: [1], [2]

Where control sits in three common research-assistant designs
DesignWho selects the next operation?ExamplePrimary check
Model callApplication or researcherExplain a band-selection ruleCheck claims and source support
Fixed workflowPredefined application logicInspect header, estimate memory, export reportCheck each stage and its inputs
Bounded agentModel within application permissionsChoose a documented inspection path after a tool resultCheck action selection, evidence and stopping rules

Give tools narrow jobs and observable results

A tool interface should make the requested operation concrete. Reading a cube header, calculating array storage and summarising a specified window are separate operations with different inputs. The result should include enough context to interpret it: dimensions, units, data type, selected bands and any error. A returned number without its assumptions is difficult to audit.

For a bounded HSI assistant, an inspect operation could return metadata without loading the cube. An estimate operation could calculate bytes from dimensions. A preview operation could produce a small image from explicitly selected bands. The application can require positive dimensions, recognised units and valid band indices before execution. These checks belong in ordinary code so that a mistaken model request cannot bypass them.

Bound an agent workflow

Enable JavaScript to explore the calculated chart.

Illustrative upper cost bound at a fixed synthetic cost of 0.2 units per call. This is not an API price or a measured performance result.

Define the research task before choosing autonomy

Consider a public demonstration task: inspect a synthetic cube, check that the wavelength list matches the band count, estimate memory, and produce a short report. This task has a clear end condition and does not require unrestricted planning. A fixed workflow may handle it well. An agent becomes more useful when it must choose among several documented inspection paths after receiving an unexpected file format or incomplete metadata.

The task brief should name the permitted inputs, output directory, available operations and resource budget. It should also define when the system must stop: an unreadable header, missing units, exhausted tool calls or a request outside the task. These are practical design choices for this example. They turn a broad instruction such as “analyse this image” into work whose completion can be checked.

Evaluate the whole system, including its failures

A research assistant needs evaluation at several levels. Check whether each tool returns the right result, whether the assistant chooses an appropriate operation, and whether its final account agrees with the recorded evidence. A correct memory calculator does not help if the assistant supplies the wrong dimensions. A correct tool call does not help if the written report changes nanometres to micrometres.

Build a small evaluation set with normal and deliberately awkward cases. Include a valid synthetic cube, a wavelength-count mismatch, absent units, a zero dimension and a tool error. Record task completion, unsupported claims, invalid requests, elapsed time and resource use separately. A single overall success percentage can hide the difference between a harmless refusal and a confident numerical error.

NIST’s Generative AI Profile is a public companion to its voluntary AI Risk Management Framework. It discusses risks across the AI lifecycle and proposes actions for managing them. For this demonstration, the useful implication is to evaluate the application in context, including its users and consequences, instead of relying solely on a model benchmark.

Sources: [3]

Treat retrieved material as data with limited authority

An assistant may read documentation, papers or metadata that contains instructions intended for a different audience. Malicious material can also try to redirect its behaviour. OWASP’s public prompt-injection guidance includes tool-parameter validation and least-privilege access. The general lesson is that a document being relevant to the question does not grant it permission to control the application.

For the HSI example, external text can inform a summary but cannot add a new upload destination or expand file access. The application should enforce the available tool set, permitted paths and resource limits. Review is appropriate before consequential actions such as publication or overwriting a curated dataset. These controls reduce specific failure paths; they are not proof that every prompt-injection attack has been solved.

Sources: [4]

Reproducibility needs a record of what actually ran

Keep a compact run record containing the task specification, input identifiers or checksums, software versions, parameters, tool requests and returned results. Record model and prompt versions when available. Distinguish the proposed operation from the executed operation. If an operation was rejected, keep the rejection reason alongside the request so a later reader can understand the missing step.

Reproducing a numerical calculation and reproducing a model’s exact wording are different goals. A deterministic tool can often repeat a calculation on the same input, while a model service may produce different text or change over time. Save the artefacts needed to verify the numerical claim. A readable explanation should point to those artefacts, rather than serve as their substitute.

A small synthetic example of enforced boundaries

The Python example below uses only the standard library and synthetic metadata. It validates dimensions, checks the wavelength count and exposes two named operations. The request to upload is rejected because no such operation exists in the registry. After two permitted calls, another call is stopped by the action budget. No image file, remote service or language model is involved.

This example demonstrates an application boundary, not a complete agent framework or security sandbox. A real system also needs appropriate process isolation, access controls and error handling. Its value here is clarity: tool availability and limits remain explicit even if the component proposing requests is later replaced by a model.

Use the assistant where its output can be checked

A useful first deployment is a repeatable inspection task with a small output and a human reader. Ask the assistant to state missing metadata, show the memory arithmetic and link the public documentation it used. Compare its report with an independently prepared answer. Expand the task only when the measured benefits justify more tool access or a longer execution budget.

Keep scientific interpretation attached to the evidence available. A preview can reveal a suspicious region, but it does not prove a material identity. A successful script can confirm that code ran, but it does not establish that the experimental design is sound. An agentic assistant earns a place in research by making these distinctions legible and by leaving work that another person can inspect.

Run the example

Prerequisite: Python 3. Examples use synthetic inputs to explain the calculation. Save the snippet as example.py and run python3 example.py.

metadata = {
    "rows": 64, "columns": 64, "bands": 6, "bytes_per_value": 4,
    "wavelengths_nm": [450, 550, 650, 750, 850, 950],
}
for key in ("rows", "columns", "bands", "bytes_per_value"):
    value = metadata[key]
    if type(value) is not int or value <= 0:
        raise ValueError(f"Invalid {key}")
if len(metadata["wavelengths_nm"]) != metadata["bands"]:
    raise ValueError("Band metadata mismatch")

def inspect():
    shape = tuple(metadata[k] for k in ("rows", "columns", "bands"))
    return f"shape={shape}; wavelengths={len(metadata['wavelengths_nm'])}"

def estimate_memory():
    size = (metadata["rows"] * metadata["columns"]
            * metadata["bands"] * metadata["bytes_per_value"])
    return f"{size / 1024:.1f} KiB"

tools = {"inspect": inspect, "estimate_memory": estimate_memory}
requests = ["inspect", "upload", "estimate_memory", "inspect"]
max_requests, max_tools = 4, 2
completed = 0
for name in requests[:max_requests]:
    if name not in tools:
        print(f"blocked: {name}")
        continue
    if completed >= max_tools:
        print("stopped: action budget")
        break
    print(f"{name}: {tools[name]()}")
    completed += 1
print(f"completed_tools: {completed}")

Verified output

inspect: shape=(64, 64, 6); wavelengths=6
blocked: upload
estimate_memory: 96.0 KiB
stopped: action budget
completed_tools: 2

Frequently asked questions

Is an agent just a larger language model?

No. Agent behaviour comes from the surrounding system: tools, state, feedback and control over the next operation. Model capability is only one component.

Does every HSI task need an agent?

No. Header inspection and fixed preprocessing steps can be implemented as ordinary workflows. Use dynamic choices where they solve a measured problem.

Can the assistant validate a scientific claim?

It can perform specified checks and organise evidence. The adequacy of those checks and the resulting scientific interpretation still require domain judgement.

What should be logged?

The task, input identifiers, software and model versions, parameters, executed tools, results and rejected requests. Avoid storing credentials or unnecessary sensitive content.

Does a strict prompt prevent prompt injection?

A prompt alone does not enforce permissions. Tool validation, restricted access and appropriate review provide additional controls, with limits that must be tested.

What is a good first evaluation?

Use a small synthetic case with a known answer, then add malformed metadata and tool failures. Assess numerical correctness and unsupported claims separately.

References and further reading

  1. Anthropic: Building effective agents
  2. ReAct research paper
  3. NIST Generative AI Profile, AI 600-1
  4. OWASP: LLM Prompt Injection Prevention