AGENTVERIFY 2027

Call for Papers

We welcome research, visions, industrial insights, and practical experiences that strengthen the foundations of trustworthy software agents.

Submission deadlines:   Abstracts are due October 18, 2026; papers are due October 23, 2026. All deadlines are Anywhere on Earth (AoE).

Contribute to AGENTVERIFY

We invite submissions that present novel research, early-stage visions, industrial insights, and practical experiences related to the engineering, evaluation, testing, debugging, benchmarking, and deployment of agentic systems for software engineering.

Submission Types

Full PapersMaximum of 8 pages, including references.
Short & Demonstration PapersMaximum of 4 pages, including references.
Extended AbstractsMaximum of 2 pages, including references.

Extended abstracts are free of APC charges and will be included in the proceedings.

Review and publication. Papers will be submitted via HotCRP and reviewed double-blind. Submissions should follow the IEEE format in line with ICSE 2027 and present an original contribution. At least one author of each accepted paper must register and present.

Accepted papers will appear in the ICSE 2027 workshop proceedings by default. A non-archival option will also be available for authors who prefer to share their work only on the workshop website.

Topics of Interest

  • Testing methodologies for LLM-based software agents.
  • Debugging, failure diagnosis, and root-cause analysis of agentic systems.
  • Harness engineering for agent execution, evaluation, and deployment.
  • Skill design, specification, validation, composition, versioning, and maintenance.
  • Benchmark construction for repository-level and tool-using software agents.
  • Reproducible evaluation protocols, grading methods, and oracle design.
  • Agent observability, telemetry, tracing, logging, and execution monitoring.
  • Sandboxing, permission control, safe tool use, and secure agent execution.
  • Evaluation under non-determinism, flaky tests, weak oracles, and environment variation.
  • Benchmark contamination, data leakage, and validity threats in agent evaluation.
  • Human-agent collaboration, oversight, intervention, and trust calibration.
  • Empirical studies, industrial experiences, tools, datasets, and open-source infrastructures.