From the Lab · Gene Avakyan

VERDICT and Explainable Autonomous Mission Replanning

A proposed framework for reviewing autonomous mission results, preserving uncertainty, and presenting feasible next-plan options for human judgment.

September 23, 2026 · by Gene Avakyan

Autonomous missions generate records, but a timeline of events is not the same as an explanation of performance. To judge an outcome, a reviewer needs the assignment as it was approved, the criteria for success, the changes made along the way, and the evidence of what happened. A single mission score can hide the step that failed. A recommendation for the next plan can be hard to trust if its assumptions and constraints remain invisible.

Conceptual VERDICT loop diagram: plan and evidence preserved, steps graded with uncertainty, options for human decision.
Conceptual VERDICT loop: preserve the plan and evidence, assess assignment steps with uncertainty, and present constrained options for a human decision.

At Edison Aerospace, I am proposing VERDICT as a framework for that interval between debrief and the next mission plan. Its purpose is to connect mission intent to observed execution, grade each assignment step against criteria declared before analysis, and present bounded options for human review. This is proposal work, not an operationally validated product or an endorsement by the Navy.

Keep intent updates and observations distinct

The original plan matters even after it changes. VERDICT would preserve that starting record and the later authorized updates, then associate incoming activity records with the relevant assignment. The system would represent dependencies: a later step may rely on an earlier one, and a missing prerequisite changes how its result should be understood. Platform records may arrive late, disagree, use different vocabularies, or omit key fields. Pretending they form one perfect history would make the debrief look more certain than it is.

The proposed approach keeps source records and their transformations traceable. Where data meanings are documented, an adapter could map them into a common review form. Where meanings are uncertain, a candidate interpretation would carry its uncertainty or remain unresolved. Competing reconstructions would remain visible when the evidence supports more than one account. A clean chart should never erase a contradiction that matters to the conclusion.

Grade assignments with visible limits

For each assignment step, VERDICT would compare the available evidence with its declared success criteria. A result could be satisfied, partial, unsatisfied, or indeterminate. The distinction between an unmet criterion and a missing observation is essential. If a required record is absent, calling the task a failure may be as misleading as calling it a success. A reviewer should also see how complete the evidence is and which dependencies affect the result.

Explanation needs the same discipline. VERDICT would label whether a statement is directly observed, an association, a model inference, a counterfactual estimate, or an operator judgment. It would show plausible alternative interpretations and limits on causal claims. A provenance trail can show where a finding came from, but provenance alone cannot prove that the finding is correct. The intended output is a reviewable argument whose uncertainty survives into the decision, not a confident narrative assembled from weak data.

Replanning is a constrained choice

After the review, a planner may ask what should change next. VERDICT would compare a preferred option, a feasible backup, and a no-change baseline when the evidence permits. Every option would state the constraints and assumptions used to judge feasibility, the expected benefit, resource implications, and sensitivity to uncertain inputs. Declared rules and deterministic checks would govern hard limits; statistical or AI methods could support uncertain reconstruction and comparison. A language model by itself would not decide that an option is feasible.

The recommendation is never meant to execute itself. The human planner would accept, modify, defer, or reject it, and that disposition would become part of the record. This preserves the value of analytic speed without hiding who made the decision. It also creates a practical way to learn from disagreements between a model’s suggestion and an experienced reviewer’s judgment.

What the proposed research must establish

Phase I is designed to select and test algorithms and produce a documented solution plan. Edison proposes a modest unclassified prototype and synthetic scenarios with incomplete, delayed, conflicting, and unfamiliar records as additional evidence. The study would examine reconstruction quality, assignment grading, uncertainty calibration, reviewability, and the time and compute required to produce results. It would also document where the system cannot reach a defensible conclusion. These evaluations are planned; no measured VERDICT result is claimed here.

The proposal contemplates commercial AI evaluation only with public or synthetic information. Later work would require locally controlled inference and an approved operating boundary before sensitive Government information could enter the system. Air-gapped deployment is a plan, not a demonstrated authorization. The broader aim is straightforward: help a human reviewer see what the mission evidence supports, what remains uncertain, and why one feasible next step may be preferable to another.

← Back to News