diff --git a/AGENTS.md b/AGENTS.md index 496cb374c86..81df141844a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,6 +31,10 @@ - Exposes tools across workflows, executions, artifacts, and metadata via `testkube mcp serve` (CLI), Docker image (`testkube/mcp-server`), or Control Plane's `/mcp` endpoint per environment. - Uses interface-based tool design; new tools need registration in both `pkg/mcp/server.go` and control plane's `mcp_handler.go`. - See `pkg/mcp/README.md` for architecture, tool patterns, and usage examples. +- Insights board tools (`pkg/mcp/tools/boards.go`) keep their rules in `pkg/mcp/boards/`: report param validation and defaults, the report-to-`/insights/*` query translation, and the layout. It is a port of the dashboard's TypeScript (`utils/insights.ts`, `DynamicFilters/types.ts`, `reports/*/type.ts` in `testkube-cloud-api/js/packages/web`), since the Control Plane stores report params opaquely. **The Control Plane's `HandlerClient` must use this package rather than reimplement it**, and a dashboard change to those files needs a matching change here; `testdata/translation_cases.json` pins the translation. +- Boards are organization-scoped and the Control Plane refuses API tokens on every board endpoint, so the board tools need a user session. `APIClient` refuses a `tkcapi_` token before sending anything and returns `tools.ErrBoardsRequireUser`. Every board write reads the board first and resends its description. Current Control Planes keep a description an update omits, but older ones clear it, so resending is what keeps it on those. +- **Board updates are optimistic-concurrency writes.** Resending a value read earlier (the description, a recomputed layout) would overwrite a concurrent edit, so every update - `update_board` and the three report tools - goes through `writeBoard` in `pkg/mcp/tools/boards.go`: it sends `expectedVersion` (the board `version` it read; every write to a board increments it), the Control Plane refuses a stale write with 409, both clients turn that into `tools.ErrBoardChanged`, and the write is rebuilt from a fresh read, up to `boardWriteAttempts` times. A write's builder must derive everything from the board it is handed, never from an earlier read. The token is a counter, not `updatedAt`: two writes can share a timestamp, and a reused token would let a stale write through. A Control Plane that predates versions returns none, and the write is then unconditional. `delete_board` is deliberately not conditional: it resends nothing it read (the read only resolves a slug to the ID it deletes by), and the Control Plane checks visibility and delete rights against the board as it is at delete time, so deleting removes the board whatever changed since the read, as deleting in the dashboard does. +- **Relative report ranges are anchored in a time zone.** The dashboard ends a `day`/`week`/`month`/`quarter` range at the viewer's local midnight, so `render_board` takes an IANA `timeZone` (default UTC) and passes it as `boards.QueryOptions.Location`; `boards` embeds `time/tzdata` because the MCP also runs from images without a zoneinfo database. ## GitOps resource sync diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 17422b1b2dc..41574708a0d 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -354,6 +354,14 @@ The Testkube CLI (`kubectl-testkube`, typically invoked as `testkube`) is a kube - Authentication tokens - Contexts (for multi-environment setups) +### MCP Server + +**Location**: [`pkg/mcp/`](pkg/mcp/) (see its [README](pkg/mcp/README.md)) + +`testkube mcp serve` and the `testkube/mcp-server` image expose Testkube to AI assistants over the Model Context Protocol. The tools in [`pkg/mcp/tools/`](pkg/mcp/tools/) depend on small client interfaces, implemented over HTTP by `APIClient` ([`pkg/mcp/api.go`](pkg/mcp/api.go)) and in-process by the Control Plane's `HandlerClient`, which registers the same tools on its per-environment `/mcp` endpoint. + +The Insights board tools keep their shared rules in [`pkg/mcp/boards/`](pkg/mcp/boards/): the report param validation and defaults, the translation of a report into the org-scoped `/insights/*` query that renders it, and the board layout. It is a port of the dashboard's rules, because the Control Plane stores report params opaquely, and both clients use it so that what the MCP writes renders in the dashboard and `render_board` returns the numbers the dashboard shows. Boards are organization-scoped and served only to user sessions, never to API tokens. Board updates are conditional on the board version the tool read (a counter every write increments), so an update built from a stale read is refused and rebuilt rather than overwriting a concurrent edit. Deleting a board is not conditional: it removes the board whatever changed since it was read. + ### External Integration: License Event Reporting The CLI reports installation lifecycle events to the Testkube license service so the diff --git a/api/executor/v1/webhook_types.go b/api/executor/v1/webhook_types.go index 8ec7e44c6b4..b94c0591f8a 100644 --- a/api/executor/v1/webhook_types.go +++ b/api/executor/v1/webhook_types.go @@ -128,7 +128,7 @@ type SecretRef struct { Key string `json:"key"` } -// +kubebuilder:validation:Enum=start-test;end-test-success;end-test-failed;end-test-aborted;end-test-timeout;become-test-up;become-test-down;become-test-failed;become-test-aborted;become-test-timeout;start-testsuite;end-testsuite-success;end-testsuite-failed;end-testsuite-aborted;end-testsuite-timeout;become-testsuite-up;become-testsuite-down;become-testsuite-failed;become-testsuite-aborted;become-testsuite-timeout;start-testworkflow;queue-testworkflow;end-testworkflow-success;end-testworkflow-failed;end-testworkflow-aborted;end-testworkflow-canceled;end-testworkflow-not-passed;become-testworkflow-up;become-testworkflow-down;become-testworkflow-failed;become-testworkflow-aborted;become-testworkflow-canceled;become-testworkflow-not-passed +// +kubebuilder:validation:Enum=start-test;end-test-success;end-test-failed;end-test-aborted;end-test-timeout;become-test-up;become-test-down;become-test-failed;become-test-aborted;become-test-timeout;start-testsuite;end-testsuite-success;end-testsuite-failed;end-testsuite-aborted;end-testsuite-timeout;become-testsuite-up;become-testsuite-down;become-testsuite-failed;become-testsuite-aborted;become-testsuite-timeout;start-testworkflow;queue-testworkflow;end-testworkflow-success;end-testworkflow-failed;end-testworkflow-aborted;end-testworkflow-canceled;end-testworkflow-not-passed;end-testworkflow-test-failure;end-testworkflow-infrastructure-failure;end-testworkflow-configuration-error;become-testworkflow-up;become-testworkflow-down;become-testworkflow-failed;become-testworkflow-aborted;become-testworkflow-canceled;become-testworkflow-not-passed type EventType string // List of EventType diff --git a/api/v1/testkube.yaml b/api/v1/testkube.yaml index 56a83adc6c8..6b360d3ba52 100644 --- a/api/v1/testkube.yaml +++ b/api/v1/testkube.yaml @@ -8029,6 +8029,9 @@ components: EventType: type: string + description: >- + The type of an event. The events end-testworkflow-test-failure, end-testworkflow-infrastructure-failure + and end-testworkflow-configuration-error select executions by the cause of the failure. enum: - queue-testworkflow - start-testworkflow @@ -8037,6 +8040,9 @@ components: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed diff --git a/cmd/api-server/main.go b/cmd/api-server/main.go index 6ac018b992b..7c1caaddf9a 100644 --- a/cmd/api-server/main.go +++ b/cmd/api-server/main.go @@ -794,7 +794,7 @@ func main() { // Push a cluster-resources snapshot to the CP on startup, on CRD informer // events, and as an hourly safety net. The CP caches it to render the // TestTrigger resourceRef picker (see AgentInventoryService). - if intconfig.ShouldPushClusterInventory(proContext) { + if intconfig.ShouldPushClusterInventory(proContext, cfg.DisableTestTriggers) { crdNotifier := inventorycontroller.StartCRDChangeNotifier(ctx, apiextClient, log.DefaultLogger) clusterResourcesController := &inventorycontroller.ClusterResourcesController{ Discoverer: api.ClusterDiscoverer, diff --git a/cmd/api-server/superagentmigration.go b/cmd/api-server/superagentmigration.go index 8e1488a5ce3..a97a4752655 100644 --- a/cmd/api-server/superagentmigration.go +++ b/cmd/api-server/superagentmigration.go @@ -43,17 +43,19 @@ type superAgentMigrationKubernetesResourceLister interface { List(ctx context.Context, list client.ObjectList, opts ...client.ListOption) error } -// skipUnownedResource reports whether a failure to sync a resource during migration is an ownership -// conflict. Those cannot be cleared by retrying, so the resource is skipped and reported; blocking -// the migration on it would wedge the agent indefinitely with no way out. -func skipUnownedResource(log superAgentMigrationLogger, kind, name string, err error) bool { - if !errors.Is(err, syncagent.ErrOwnershipConflict) { +// skipRejectedResource reports whether the migration must skip a resource that failed to sync. +// It skips and logs a rejection that no retry clears (see syncagent.IsRejection). Without the skip, +// the migration retries that resource forever and never completes. +func skipRejectedResource(log superAgentMigrationLogger, kind, name string, err error) bool { + if !syncagent.IsRejection(err) { return false } - log.Errorw("resource is owned by another GitOps agent, skipping it during SuperAgent migration. It will not be present in the Control Plane until its ownership is resolved.", - kind, name, - "error", err.Error()) + msg := "resource is owned by another GitOps agent, skipping it during SuperAgent migration. It will not be present in the Control Plane until its ownership is resolved." + if errors.Is(err, syncagent.ErrInvalidResource) { + msg = "the Control Plane rejected the resource as invalid, skipping it during SuperAgent migration. It will not be present in the Control Plane until the resource is fixed." + } + log.Errorw(msg, kind, name, "error", err.Error()) return true } @@ -181,7 +183,7 @@ func migrateSuperAgent(ctx context.Context, log superAgentMigrationLogger, cfg s for _, t := range testTriggerList.Items { for { if err := syncStore.UpdateOrCreateTestTrigger(ctx, t); err != nil { - if skipUnownedResource(log, "TestTrigger", t.Name, err) { + if skipRejectedResource(log, "TestTrigger", t.Name, err) { break } retryAfter := b.Duration() @@ -201,7 +203,7 @@ func migrateSuperAgent(ctx context.Context, log superAgentMigrationLogger, cfg s for _, t := range testWorkflowTemplateList.Items { for { if err := syncStore.UpdateOrCreateTestWorkflowTemplate(ctx, t); err != nil { - if skipUnownedResource(log, "TestWorkflowTemplate", t.Name, err) { + if skipRejectedResource(log, "TestWorkflowTemplate", t.Name, err) { break } retryAfter := b.Duration() @@ -219,7 +221,7 @@ func migrateSuperAgent(ctx context.Context, log superAgentMigrationLogger, cfg s for _, t := range testWorkflowList.Items { for { if err := syncStore.UpdateOrCreateTestWorkflow(ctx, t); err != nil { - if skipUnownedResource(log, "TestWorkflow", t.Name, err) { + if skipRejectedResource(log, "TestWorkflow", t.Name, err) { break } retryAfter := b.Duration() @@ -237,7 +239,7 @@ func migrateSuperAgent(ctx context.Context, log superAgentMigrationLogger, cfg s for _, t := range webhookList.Items { for { if err := syncStore.UpdateOrCreateWebhook(ctx, t); err != nil { - if skipUnownedResource(log, "Webhook", t.Name, err) { + if skipRejectedResource(log, "Webhook", t.Name, err) { break } retryAfter := b.Duration() @@ -255,7 +257,7 @@ func migrateSuperAgent(ctx context.Context, log superAgentMigrationLogger, cfg s for _, t := range webhookTemplateList.Items { for { if err := syncStore.UpdateOrCreateWebhookTemplate(ctx, t); err != nil { - if skipUnownedResource(log, "WebhookTemplate", t.Name, err) { + if skipRejectedResource(log, "WebhookTemplate", t.Name, err) { break } retryAfter := b.Duration() diff --git a/cmd/kubectl-testkube/commands/agent.go b/cmd/kubectl-testkube/commands/agent.go index 27ba04a7d4d..6d75b646026 100644 --- a/cmd/kubectl-testkube/commands/agent.go +++ b/cmd/kubectl-testkube/commands/agent.go @@ -11,8 +11,9 @@ import ( func NewAgentCmd() *cobra.Command { cmd := &cobra.Command{ - Use: "agent", - Short: "Testkube Pro Agent related commands", + Use: "runner", + Aliases: []string{"agent"}, + Short: "Testkube Pro Runner related commands", Run: func(cmd *cobra.Command, args []string) { client, _, err := common.GetClient(cmd) ui.ExitOnError("getting client", err) diff --git a/cmd/kubectl-testkube/commands/agent/debug.go b/cmd/kubectl-testkube/commands/agent/debug.go index 23c15d2efef..f66423205bc 100644 --- a/cmd/kubectl-testkube/commands/agent/debug.go +++ b/cmd/kubectl-testkube/commands/agent/debug.go @@ -7,8 +7,8 @@ import ( func NewDebugAgentCmd() *cobra.Command { cmd := &cobra.Command{ Use: "debug", - Short: "Debug Agent info", - Deprecated: "use `testkube debug agent` instead", + Short: "Debug Runner info", + Deprecated: "use `testkube debug runner` instead", } return cmd diff --git a/cmd/kubectl-testkube/commands/agent/migrate.go b/cmd/kubectl-testkube/commands/agent/migrate.go index 054fb888144..6cba34508d9 100644 --- a/cmd/kubectl-testkube/commands/agent/migrate.go +++ b/cmd/kubectl-testkube/commands/agent/migrate.go @@ -6,9 +6,10 @@ import ( func NewMigrateAgentCmd() *cobra.Command { cmd := &cobra.Command{ - Use: "agent", - Short: "manual migrate agent command", - Long: `migrate agent command will run agent migrations greater or equals current version`, + Use: "runner", + Aliases: []string{"agent"}, + Short: "manual migrate runner command", + Long: `migrate runner command will run runner migrations greater or equals current version`, Run: func(cmd *cobra.Command, args []string) { // TODO: Delete, as we don't have any migrations }, diff --git a/cmd/kubectl-testkube/commands/agents/create.go b/cmd/kubectl-testkube/commands/agents/create.go index 02a937b0ec9..fc6ab965854 100644 --- a/cmd/kubectl-testkube/commands/agents/create.go +++ b/cmd/kubectl-testkube/commands/agents/create.go @@ -19,8 +19,9 @@ func NewCreateAgentCommand() *cobra.Command { agentType string ) cmd := &cobra.Command{ - Use: "agent", - Args: cobra.ExactArgs(1), + Use: "runner", + Aliases: []string{"agent"}, + Args: cobra.ExactArgs(1), Run: func(cmd *cobra.Command, args []string) { // Check for deprecated --type flag usage if cmd.Flags().Changed("type") { @@ -62,8 +63,8 @@ func NewCreateAgentCommand() *cobra.Command { enableWebhooks, ) ui.NL() - ui.Hint("Install the agent with command:") - installCmd := fmt.Sprintf("testkube install agent %s --secret %s", agent.Name, agent.SecretKey) + ui.Hint("Install the runner with command:") + installCmd := fmt.Sprintf("testkube install runner %s --secret %s", agent.Name, agent.SecretKey) if enableExecution { installCmd += " --execution" } @@ -80,11 +81,11 @@ func NewCreateAgentCommand() *cobra.Command { }, } - cmd.Flags().StringSliceVarP(&environmentIds, "env", "e", nil, "environment ID or slug that the agent have access to") + cmd.Flags().StringSliceVarP(&environmentIds, "env", "e", nil, "environment ID or slug that the runner have access to") cmd.Flags().StringSliceVarP(&labelPairs, "label", "l", nil, "label key value pair: --label key1=value1") - cmd.Flags().BoolVar(&global, "global", false, "make it global agent") - cmd.Flags().StringVar(&group, "group", "", "make it grouped agent") - cmd.Flags().BoolVar(&floating, "floating", false, "create as a floating agent") + cmd.Flags().BoolVar(&global, "global", false, "make it global runner") + cmd.Flags().StringVar(&group, "group", "", "make it grouped runner") + cmd.Flags().BoolVar(&floating, "floating", false, "create as a floating runner") // Components selection common.AddExecutionCapabilityFlags(cmd) @@ -93,7 +94,7 @@ func NewCreateAgentCommand() *cobra.Command { cmd.Flags().Bool("webhooks", false, "enable webhooks capability") // Deprecated flag - cmd.Flags().StringVarP(&agentType, "type", "t", "", "[DEPRECATED] agent type - use capability flags instead") + cmd.Flags().StringVarP(&agentType, "type", "t", "", "[DEPRECATED] runner type - use capability flags instead") cmd.Flags().MarkDeprecated("type", "use --execution, --listener, --gitops, and/or --webhooks instead") return cmd diff --git a/cmd/kubectl-testkube/commands/agents/delete.go b/cmd/kubectl-testkube/commands/agents/delete.go index 2084ee4893a..9fe665c551f 100644 --- a/cmd/kubectl-testkube/commands/agents/delete.go +++ b/cmd/kubectl-testkube/commands/agents/delete.go @@ -1,6 +1,7 @@ package agents import ( + "errors" "fmt" "os" "strings" @@ -17,24 +18,24 @@ func NewDeleteAgentCommand() *cobra.Command { deleteAgent, noDeleteAgent bool ) cmd := &cobra.Command{ - Use: "agent", - Aliases: []string{"runner"}, + Use: "runner", + Aliases: []string{"agent"}, Args: cobra.ExactArgs(1), Run: func(cmd *cobra.Command, args []string) { if !uninstall && !noUninstall { - uninstall = ui.Confirm("should it uninstall agent?") + uninstall = ui.Confirm("should it uninstall runner?") } if !deleteAgent && !noDeleteAgent { - deleteAgent = ui.Confirm("should it delete agent in the Control Plane?") + deleteAgent = ui.Confirm("should it delete runner in the Control Plane?") } UiDeleteAgent(cmd, args[0], uninstall, deleteAgent) }, } - cmd.Flags().BoolVarP(&uninstall, "uninstall", "u", false, "should it uninstall the agent too") - cmd.Flags().BoolVarP(&noUninstall, "no-uninstall", "U", false, "should it keep the agent installed") - cmd.Flags().BoolVarP(&deleteAgent, "delete", "d", false, "should it delete agent in the Control Plane") - cmd.Flags().BoolVarP(&noDeleteAgent, "no-delete", "D", false, "should it keep the agent definition in the Control Plane") + cmd.Flags().BoolVarP(&uninstall, "uninstall", "u", false, "should it uninstall the runner too") + cmd.Flags().BoolVarP(&noUninstall, "no-uninstall", "U", false, "should it keep the runner installed") + cmd.Flags().BoolVarP(&deleteAgent, "delete", "d", false, "should it delete runner in the Control Plane") + cmd.Flags().BoolVarP(&noDeleteAgent, "no-delete", "D", false, "should it keep the runner definition in the Control Plane") return cmd } @@ -64,13 +65,23 @@ func UiUninstallCRD(cmd *cobra.Command) { spinner := ui.NewSpinner("Fetching current CRDs") currentNamespace, currentReleaseName, installed, err := GetCRDInstallation() if err != nil { - spinner.Fail(err) - os.Exit(1) + spinner.Fail() + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceLookupFailed, + "Error getting the installed CRDs", + common2.ClusterLookupHint, + err, + )) } if installed && currentReleaseName == "" { - spinner.Fail("The CRDs are installed, but they are not managed by our Helm Chart") - os.Exit(1) + spinner.Fail() + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidInstallConfig, + "The CRDs are not managed by the Testkube Helm Chart", + "Delete the Testkube CRDs by hand, or install them with `testkube install crd` so that Helm owns them", + errors.New("the CRDs are installed, but they carry no Helm release annotation"), + )) } if installed { @@ -81,30 +92,47 @@ func UiUninstallCRD(cmd *cobra.Command) { } spinner = ui.NewSpinner("Uninstalling CRDs") - cliErr := common2.HelmUninstall(currentNamespace, currentReleaseName) - if cliErr != nil { - cliErr.Print() - os.Exit(1) - } + common2.HandleCLIError(common2.HelmUninstall(currentNamespace, currentReleaseName)) spinner.Success() } func UiDeleteAgent(cmd *cobra.Command, name string, uninstall, deleteAgent bool) { agent, err := GetControlPlaneAgent(cmd, name) - ui.ExitOnError("getting agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerGetFailed, + "Error getting the runner", + common2.RunnerLookupHint, + err, + )) + } - // Uninstall the Agent + // Uninstall the Runner if uninstall { var nses []string if agent.Namespace != "" { nses = append(nses, agent.Namespace) } else { nses, err = GetKubernetesNamespaces() - ui.ExitOnError("getting namespaces", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceLookupFailed, + "Error listing the Kubernetes namespaces", + common2.ClusterLookupHint, + err, + )) + } } agents, err := GetKubernetesAgents(nses) - ui.ExitOnError("getting agents", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceLookupFailed, + "Error getting the runners running in the cluster", + common2.ClusterLookupHint, + err, + )) + } var kubernetesAgent *internalAgent for i := range agents { @@ -114,24 +142,32 @@ func UiDeleteAgent(cmd *cobra.Command, name string, uninstall, deleteAgent bool) } } if kubernetesAgent == nil { - ui.Failf("kubernetes agent not found: namespaces: %s", strings.Join(nses, ", ")) + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceNotFound, + "Runner not installed in the cluster", + "Pass '--no-uninstall' to delete the runner in the Control Plane only, or check that your kubeconfig points at the cluster the runner runs in", + fmt.Errorf("kubernetes runner not found: namespaces: %s", strings.Join(nses, ", ")), + )) return } spinner := ui.NewSpinner("Running Helm command...") - cliErr := common2.HelmUninstall(kubernetesAgent.Pod.Namespace, fmt.Sprintf("testkube-%s", agent.Name)) - if cliErr != nil { - cliErr.Print() - os.Exit(1) - } + common2.HandleCLIError(common2.HelmUninstall(kubernetesAgent.Pod.Namespace, fmt.Sprintf("testkube-%s", agent.Name))) spinner.Success() } // Delete the Agent if deleteAgent { - spinner := ui.NewSpinner("Deleting agent in the Control Plane...") + spinner := ui.NewSpinner("Deleting runner in the Control Plane...") err := DeleteControlPlaneAgent(cmd, agent.ID) - ui.ExitOnError("deleting agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerWriteFailed, + "Error deleting the runner", + common2.RunnerWriteHint, + err, + )) + } spinner.Success() } } diff --git a/cmd/kubectl-testkube/commands/agents/enable.go b/cmd/kubectl-testkube/commands/agents/enable.go index 7f16ed36f8a..a6bd0e582ad 100644 --- a/cmd/kubectl-testkube/commands/agents/enable.go +++ b/cmd/kubectl-testkube/commands/agents/enable.go @@ -5,6 +5,7 @@ import ( "github.com/spf13/cobra" + common2 "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/common" "github.com/kubeshop/testkube/internal/common" "github.com/kubeshop/testkube/pkg/cloud/client" "github.com/kubeshop/testkube/pkg/ui" @@ -12,8 +13,8 @@ import ( func NewEnableAgentCommand() *cobra.Command { cmd := &cobra.Command{ - Use: "agent ", - Aliases: []string{"runner", "gitops"}, + Use: "runner ", + Aliases: []string{"agent", "gitops"}, Args: cobra.ExactArgs(1), Run: func(cmd *cobra.Command, args []string) { UiEnableAgent(cmd, strings.Join(args, "")) @@ -25,8 +26,8 @@ func NewEnableAgentCommand() *cobra.Command { func NewDisableAgentCommand() *cobra.Command { cmd := &cobra.Command{ - Use: "agent ", - Aliases: []string{"runner", "gitops"}, + Use: "runner ", + Aliases: []string{"agent", "gitops"}, Args: cobra.ExactArgs(1), Run: func(cmd *cobra.Command, args []string) { UiDisableAgent(cmd, strings.Join(args, "")) @@ -38,15 +39,29 @@ func NewDisableAgentCommand() *cobra.Command { func UiEnableAgent(cmd *cobra.Command, name string) { agent, err := GetControlPlaneAgent(cmd, name) - ui.ExitOnError("getting agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerGetFailed, + "Error getting the runner", + common2.RunnerLookupHint, + err, + )) + } if agent.Disabled { agent, err = UpdateAgent(cmd, agent.ID, client.AgentInput{ Disabled: common.Ptr(false), }) - ui.ExitOnError("updating agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerWriteFailed, + "Error enabling the runner", + common2.RunnerWriteHint, + err, + )) + } } else { - ui.Print("Agent is already enabled.") + ui.Print("Runner is already enabled.") } PrintControlPlaneAgent(*agent) @@ -54,15 +69,29 @@ func UiEnableAgent(cmd *cobra.Command, name string) { func UiDisableAgent(cmd *cobra.Command, name string) { agent, err := GetControlPlaneAgent(cmd, name) - ui.ExitOnError("getting agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerGetFailed, + "Error getting the runner", + common2.RunnerLookupHint, + err, + )) + } if !agent.Disabled { agent, err = UpdateAgent(cmd, agent.ID, client.AgentInput{ Disabled: common.Ptr(true), }) - ui.ExitOnError("updating agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerWriteFailed, + "Error disabling the runner", + common2.RunnerWriteHint, + err, + )) + } } else { - ui.Print("Agent is already disabled.") + ui.Print("Runner is already disabled.") } PrintControlPlaneAgent(*agent) diff --git a/cmd/kubectl-testkube/commands/agents/get.go b/cmd/kubectl-testkube/commands/agents/get.go index 20affd2a0ba..0672a48fea2 100644 --- a/cmd/kubectl-testkube/commands/agents/get.go +++ b/cmd/kubectl-testkube/commands/agents/get.go @@ -6,6 +6,7 @@ import ( "github.com/spf13/cobra" + "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/common" "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/common/render" "github.com/kubeshop/testkube/cmd/kubectl-testkube/config" "github.com/kubeshop/testkube/pkg/ui" @@ -20,13 +21,13 @@ func NewGetAgentCommand() *cobra.Command { ) cmd := &cobra.Command{ Args: cobra.MaximumNArgs(1), - Use: "agent [name]", - Short: "Get agents registered in the current environment", - Long: `Get details of a specific agent or list all agents. By default only active agents in the current environment are shown. Use --all-environments to list across environments, --show-deleted to view deleted agents, or --show-unknown to find cluster agents not registered in the control plane.`, - Aliases: []string{"agents", "a"}, + Use: "runner [name]", + Short: "Get runners registered in the current environment", + Long: `Get details of a specific runner or list all runners. By default only active runners in the current environment are shown. Use --all-environments to list across environments, --show-deleted to view deleted runners, or --show-unknown to find cluster runners not registered in the control plane.`, + Aliases: []string{"runners", "agent", "agents", "a"}, PreRun: func(cmd *cobra.Command, args []string) { if allEnvironments && showUnknown { - ui.Warn("Note: --all-environments is ignored when using --show-unknown (unknown agents have no environment registration)") + ui.Warn("Note: --all-environments is ignored when using --show-unknown (unknown runners have no environment registration)") allEnvironments = false } }, @@ -40,22 +41,43 @@ func NewGetAgentCommand() *cobra.Command { } cmd.Flags().BoolVar(&decryptSecretKey, "decrypted-secret", false, "should it fetch decrypted secret key") - cmd.Flags().BoolVar(&showUnknown, "show-unknown", false, "show only unknown agents (agents in cluster not registered in control plane)") - cmd.Flags().BoolVar(&showDeleted, "show-deleted", false, "show only deleted agents") - cmd.Flags().BoolVar(&allEnvironments, "all-environments", false, "show agents from all environments (not just current environment)") + cmd.Flags().BoolVar(&showUnknown, "show-unknown", false, "show only unknown runners (runners in cluster not registered in control plane)") + cmd.Flags().BoolVar(&showDeleted, "show-deleted", false, "show only deleted runners") + cmd.Flags().BoolVar(&allEnvironments, "all-environments", false, "show runners from all environments (not just current environment)") return cmd } func UiGetAgent(cmd *cobra.Command, agentId string, decryptSecretKey bool) { registeredAgents, err := GetControlPlaneAgents(cmd, true) - ui.ExitOnError("getting agents", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrRunnerGetFailed, + "Error getting the runners", + common.RunnerLookupHint, + err, + )) + } namespaces, err := GetKubernetesNamespaces() - ui.ExitOnError("listing namespaces", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrResourceLookupFailed, + "Error listing the Kubernetes namespaces", + common.ClusterLookupHint, + err, + )) + } agents, err := GetKubernetesAgents(namespaces) - ui.ExitOnError("listing pods", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrResourceLookupFailed, + "Error getting the runners running in the cluster", + common.ClusterLookupHint, + err, + )) + } agents = CombineAgents(agents, registeredAgents) @@ -67,12 +89,24 @@ func UiGetAgent(cmd *cobra.Command, agentId string, decryptSecretKey bool) { } } if agent == nil { - ui.Fail(fmt.Errorf("agent '%s' not found", agentId)) + common.HandleCLIError(common.NewCLIError( + common.TKErrResourceNotFound, + "Runner not found", + "Check the runner name or ID, or list the runners with `testkube get runners`. Add --show-unknown to include cluster runners that are not registered, and --show-deleted to include deleted ones", + fmt.Errorf("runner '%s' not found", agentId), + )) } if decryptSecretKey { secretKey, err := GetControlPlaneAgentSecretKey(cmd, agent.Registered.ID) - ui.ExitOnError("failed to decrypt secret key", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrRunnerGetFailed, + "Error getting the decrypted runner secret key", + "Check that your credentials are valid and that your user can read the secret key of this runner", + err, + )) + } agent.Registered.SecretKey = secretKey } @@ -81,17 +115,39 @@ func UiGetAgent(cmd *cobra.Command, agentId string, decryptSecretKey bool) { func UiListAgents(cmd *cobra.Command, showUnknown bool, showDeleted bool, allEnvironments bool) { registeredAgents, err := GetControlPlaneAgents(cmd, showDeleted) - ui.ExitOnError("getting agents", err) + if err != nil { + // The hint of a lookup by name points at `testkube get runners`, which is this command. + common.HandleCLIError(common.NewCLIError( + common.TKErrRunnerGetFailed, + "Error getting the runners", + "Check that your credentials are valid and that the current context points at the organization and environment you expect", + err, + )) + } // Filter agents by current environment (matching dashboard behavior) unless --all-environments is set if !allEnvironments { cfg, err := config.Load() - ui.ExitOnError("loading config", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrConfigInitFailed, + "Error loading testkube config file", + common.ConfigFileHint, + err, + )) + } registeredAgents = FilterAgentsByEnvironment(registeredAgents, cfg.CloudContext.EnvironmentId) } agents, err := GetKubernetesAgents([]string{""}) - ui.ExitOnError("listing pods", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrResourceLookupFailed, + "Error getting the runners running in the cluster", + common.ClusterLookupHint, + err, + )) + } agents = CombineAgents(agents, registeredAgents) @@ -130,7 +186,7 @@ func UiListAgents(cmd *cobra.Command, showUnknown bool, showDeleted bool, allEnv agents = filteredAgents if len(agents) == 0 { - ui.Print(ui.LightGray("\nNo agents found")) + ui.Print(ui.LightGray("\nNo runners found")) return } diff --git a/cmd/kubectl-testkube/commands/agents/install.go b/cmd/kubectl-testkube/commands/agents/install.go index 4125efaf9ed..750c88e038a 100644 --- a/cmd/kubectl-testkube/commands/agents/install.go +++ b/cmd/kubectl-testkube/commands/agents/install.go @@ -22,8 +22,9 @@ func NewInstallAgentCommand() *cobra.Command { var namespace string cmd := &cobra.Command{ - Use: "agent ", - Args: cobra.MaximumNArgs(1), + Use: "runner ", + Aliases: []string{"agent"}, + Args: cobra.MaximumNArgs(1), Run: func(cmd *cobra.Command, args []string) { // Check for deprecated --type flag usage if cmd.Flags().Changed("type") { @@ -41,7 +42,7 @@ func NewInstallAgentCommand() *cobra.Command { }, } - cmd.Flags().StringVarP(&namespace, "namespace", "n", "", "namespace to install the agent") + cmd.Flags().StringVarP(&namespace, "namespace", "n", "", "namespace to install the runner") common2.PopulateRunnerFlags(cmd) return cmd } @@ -72,13 +73,23 @@ func UiInstallCRD(cmd *cobra.Command, namespace string, releaseName string, dryR spinner := ui.NewSpinner("Fetching current CRDs") currentNamespace, currentReleaseName, installed, err := GetCRDInstallation() if err != nil { - spinner.Fail(err) - os.Exit(1) + spinner.Fail() + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceLookupFailed, + "Error getting the installed CRDs", + common2.ClusterLookupHint, + err, + )) } if installed && currentReleaseName == "" { - spinner.Fail("The CRDs are installed, but they are not managed by our Helm Chart") - os.Exit(1) + spinner.Fail() + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidInstallConfig, + "The CRDs are not managed by the Testkube Helm Chart", + "Delete the Testkube CRDs by hand, so that this command can install them with Helm", + fmt.Errorf("the CRDs are installed, but they carry no Helm release annotation"), + )) } if installed { @@ -92,11 +103,7 @@ func UiInstallCRD(cmd *cobra.Command, namespace string, releaseName string, dryR } opts := CreateCRDsHelmOptions(namespace, releaseName, dryRun, nil) - cliErr := common2.HelmUpgradeOrInstallGeneric(opts) - if cliErr != nil { - cliErr.Print() - os.Exit(1) - } + common2.HandleCLIError(common2.HelmUpgradeOrInstallGeneric(opts)) spinner.Success("CRDs installed") } @@ -128,16 +135,36 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext if globalTemplatePath != "" { var err error globalTemplate, err = os.ReadFile(globalTemplatePath) - ui.ExitOnError("reading global template", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "Error reading the global template", + "Check that the '--global-template-path' value points at a readable file", + err, + )) + } globalTemplateMap := make(map[string]interface{}) err = yaml.Unmarshal(globalTemplate, &globalTemplateMap) - ui.ExitOnError("reading global template", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "Error parsing the global template", + "The file that '--global-template-path' names must be a YAML Test Workflow template", + err, + )) + } if spec, ok := globalTemplateMap["spec"]; ok { globalTemplate, err = json.Marshal(spec) - ui.ExitOnError("marshalling global template", err) } else { globalTemplate, err = json.Marshal(globalTemplateMap) - ui.ExitOnError("marshalling global template", err) + } + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "Error converting the global template", + "The file that '--global-template-path' names must hold values that convert to JSON", + err, + )) } } @@ -146,8 +173,13 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext if name != "" { var err error agent, err = GetControlPlaneAgent(cmd, name) - if !autoCreate { - ui.ExitOnError("getting agent", err) + if err != nil && !autoCreate { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerGetFailed, + "Error getting the runner", + "Check the runner name or ID and that your credentials are valid, or pass '--create' to create the runner", + err, + )) } if agent != nil { PrintControlPlaneAgent(*agent) @@ -178,14 +210,26 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext // Load agents from the Control Plane and select one if agent == nil { agents, err := GetControlPlaneAgents(cmd, false) - ui.ExitOnError("listing agents", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerGetFailed, + "Error getting the runners", + common2.RunnerLookupHint, + err, + )) + } if name == "" { - name = ui.Select("select agent", common.MapSlice(agents, func(t cloudclient.Agent) string { + name = ui.Select("select runner", common.MapSlice(agents, func(t cloudclient.Agent) string { return t.Name })) if name == "" { - ui.Failf("agent name not provided") + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "No runner name provided", + "Pass the runner name as an argument, for example `testkube install runner my-runner`", + fmt.Errorf("runner name not provided"), + )) } } @@ -199,7 +243,12 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext // Fail if there is no matching agent available if agent == nil { - ui.Failf("agent %s not found", name) + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceNotFound, + "Runner not found", + "Check the runner name or ID and list the runners with `testkube get runners`, or pass '--create' to create it", + fmt.Errorf("runner %s not found", name), + )) return } @@ -209,7 +258,14 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext if agent.SecretKey == "" { secretKey, err := GetControlPlaneAgentSecretKey(cmd, agent.ID) - ui.ExitOnError("failed to fetch the secret key", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerGetFailed, + "Error getting the runner secret key", + "Check that your credentials are valid and that your user can read the secret key of this runner, or pass it with '--secret'", + err, + )) + } agent.SecretKey = secretKey } @@ -240,13 +296,25 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext } ns = ui.TextInput("namespace to install", defaultNs) if ns == "" { - ui.Failf("you need to select namespace to install") + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "No namespace provided", + "Pass the namespace with '--namespace', or type one at the prompt", + fmt.Errorf("you need to select namespace to install"), + )) } } // Load the Cloud settings cfg, err := config.Load() - ui.ExitOnError("loading config file", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrConfigInitFailed, + "Error loading testkube config file", + common2.ConfigFileHint, + err, + )) + } skipTLS := common2.ResolveSkipTLS(cmd, &cfg) opts := &common2.HelmOptions{} common2.ProcessMasterFlags(cmd, opts, &cfg) @@ -297,11 +365,7 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext helmOpts.Values["globalTemplate.inline"] = true helmOpts.Values["globalTemplate.spec"] = string(globalTemplate) } - cliErr := common2.HelmUpgradeOrInstallGeneric(helmOpts) - if cliErr != nil { - cliErr.Print() - os.Exit(1) - } + common2.HandleCLIError(common2.HelmUpgradeOrInstallGeneric(helmOpts)) if dryRun { return @@ -310,7 +374,14 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext spinner.Success() agents, err := GetKubernetesAgents([]string{ns}) - ui.ExitOnError("getting agents in kubernetes", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceLookupFailed, + "Error getting the runners running in the cluster", + common2.ClusterLookupHint, + err, + )) + } var foundAgent *internalAgent for i := range agents { @@ -321,7 +392,12 @@ func UiInstallAgent(cmd *cobra.Command, name string, defaultLabels []string, ext } if foundAgent == nil { - ui.Failf("not found the agent installed in namespace '%s'", ns) + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrResourceNotFound, + "Runner not found in the cluster", + "The Helm release installed, but no runner Pod carries its id yet. Check the Pods of the namespace, for example with `kubectl get pods -n `", + fmt.Errorf("not found the runner installed in namespace '%s'", ns), + )) return } diff --git a/cmd/kubectl-testkube/commands/agents/rotate_key.go b/cmd/kubectl-testkube/commands/agents/rotate_key.go index 7821f9faf73..8940159b27a 100644 --- a/cmd/kubectl-testkube/commands/agents/rotate_key.go +++ b/cmd/kubectl-testkube/commands/agents/rotate_key.go @@ -20,12 +20,19 @@ func NewRotateKeyCommand() *cobra.Command { cmd := &cobra.Command{ Use: "rotate-key ", - Short: "Rotate the secret key for an agent", - Long: "Rotate the secret key for an agent with a configurable grace period during which the old key remains valid", + Short: "Rotate the secret key for a runner", + Long: "Rotate the secret key for a runner with a configurable grace period during which the old key remains valid", Args: cobra.ExactArgs(1), PersistentPreRun: func(cmd *cobra.Command, args []string) { cfg, err := config.Load() - ui.ExitOnError("loading config", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrConfigInitFailed, + "Error loading testkube config file", + common.ConfigFileHint, + err, + )) + } common.UiContextHeader(cmd, cfg) validator.PersistentPreRunVersionCheck(cmd, common.Version) }, @@ -36,9 +43,9 @@ func NewRotateKeyCommand() *cobra.Command { agent, err := GetControlPlaneAgent(cmd, nameOrID) if err != nil { common.HandleCLIError(common.NewCLIError( - common.TKErrAgentGetFailed, - "Failed to get agent", - "Verify the agent name or ID is correct and your credentials are valid", + common.TKErrRunnerGetFailed, + "Error getting the runner", + common.RunnerLookupHint, err, )) return @@ -46,7 +53,7 @@ func NewRotateKeyCommand() *cobra.Command { // Confirm unless --yes if !yes { - ok := ui.Confirm(fmt.Sprintf("Rotate secret key for agent '%s'?", agent.Name)) + ok := ui.Confirm(fmt.Sprintf("Rotate secret key for runner '%s'?", agent.Name)) if !ok { return } @@ -56,9 +63,9 @@ func NewRotateKeyCommand() *cobra.Command { result, err := RotateControlPlaneAgentKey(cmd, agent.ID, gracePeriod) if err != nil { common.HandleCLIError(common.NewCLIError( - common.TKErrAgentRotateKeyFailed, - "Failed to rotate agent secret key", - "Verify the agent exists and your credentials are valid", + common.TKErrRunnerRotateKeyFailed, + "Error rotating the runner secret key", + "Check that your credentials are valid and that the '--grace-period' value is one the control plane accepts, for example 24h or 0s", err, )) return @@ -67,7 +74,7 @@ func NewRotateKeyCommand() *cobra.Command { // Display results ui.Success("Secret key rotated successfully") fmt.Println() - ui.Warn("Agent: ", agent.Name) + ui.Warn("Runner: ", agent.Name) ui.Warn("New Secret Key:", result.SecretKey) if result.GracePeriod != "" { ui.Warn("Grace Period: ", result.GracePeriod) @@ -79,7 +86,7 @@ func NewRotateKeyCommand() *cobra.Command { } fmt.Println() - ui.Info("To update the agent's Kubernetes secret, run:") + ui.Info("To update the runner's Kubernetes secret, run:") fmt.Printf(" kubectl create secret generic testkube-agent-secret --from-literal=TESTKUBE_PRO_API_KEY=%s --dry-run=client -o yaml | kubectl apply -f -\n", result.SecretKey) fmt.Println() }, diff --git a/cmd/kubectl-testkube/commands/agents/rotate_registration_token.go b/cmd/kubectl-testkube/commands/agents/rotate_registration_token.go index 5d5ddd2c143..e03ddd659d4 100644 --- a/cmd/kubectl-testkube/commands/agents/rotate_registration_token.go +++ b/cmd/kubectl-testkube/commands/agents/rotate_registration_token.go @@ -1,6 +1,7 @@ package agents import ( + "errors" "fmt" "time" @@ -23,18 +24,40 @@ func NewRotateRegistrationTokenCommand() *cobra.Command { Args: cobra.MaximumNArgs(1), PersistentPreRun: func(cmd *cobra.Command, args []string) { cfg, err := config.Load() - ui.ExitOnError("loading config", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrConfigInitFailed, + "Error loading testkube config file", + common.ConfigFileHint, + err, + )) + } common.UiContextHeader(cmd, cfg) validator.PersistentPreRunVersionCheck(cmd, common.Version) }, Run: func(cmd *cobra.Command, args []string) { cfg, err := config.Load() - ui.ExitOnError("loading config", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrConfigInitFailed, + "Error loading testkube config file", + common.ConfigFileHint, + err, + )) + } envID := cfg.CloudContext.EnvironmentId if len(args) == 1 { envID = args[0] } + if envID == "" { + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "No environment selected", + "Pass the environment id as an argument, or select one with `testkube set context --env-id `", + errors.New("no environment id given and none set in the current context"), + )) + } if !yes && !ui.Confirm(fmt.Sprintf("Rotate registration token for environment ID '%s'?", envID)) { return @@ -44,9 +67,9 @@ func NewRotateRegistrationTokenCommand() *cobra.Command { result, err := client.RotateRegistrationToken(cmd.Context(), envID, gracePeriod) if err != nil { common.HandleCLIError(common.NewCLIError( - common.TKErrAgentRotateRegistrationTokenFailed, - "Failed to rotate environment registration token", - "Verify the environment ID is correct, your credentials are valid, and your user is an organization admin or owner.", + common.TKErrRunnerRotateRegistrationTokenFailed, + "Error rotating the environment registration token", + "Verify the environment ID is correct, your credentials are valid, and your user is an organization admin or owner", err, )) return diff --git a/cmd/kubectl-testkube/commands/agents/update.go b/cmd/kubectl-testkube/commands/agents/update.go index 847b32cf8ab..3726e50c9d6 100644 --- a/cmd/kubectl-testkube/commands/agents/update.go +++ b/cmd/kubectl-testkube/commands/agents/update.go @@ -6,9 +6,9 @@ import ( "github.com/spf13/cobra" + common2 "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/common" "github.com/kubeshop/testkube/internal/common" "github.com/kubeshop/testkube/pkg/cloud/client" - "github.com/kubeshop/testkube/pkg/ui" ) func NewUpdateAgentCommand() *cobra.Command { @@ -20,8 +20,8 @@ func NewUpdateAgentCommand() *cobra.Command { ) cmd := &cobra.Command{ - Use: "agent ", - Aliases: []string{"runner"}, + Use: "runner ", + Aliases: []string{"agent"}, Args: cobra.ExactArgs(1), Run: func(cmd *cobra.Command, args []string) { UiUpdateAgent(cmd, strings.Join(args, ""), setLabels, deleteLabels, runnerMode, groupName) @@ -38,13 +38,34 @@ func NewUpdateAgentCommand() *cobra.Command { func UiUpdateAgent(cmd *cobra.Command, name string, setLabels, deleteLabels []string, runnerMode string, groupName string) { agent, err := GetControlPlaneAgent(cmd, name) - ui.ExitOnError("getting agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerGetFailed, + "Error getting the runner", + common2.RunnerLookupHint, + err, + )) + } input, err := buildUpdateAgentInput(agent, setLabels, deleteLabels, runnerMode, groupName) - ui.ExitOnError("preparing agent update", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "Error preparing the runner update", + "Check the '--runner-mode', '--group-name', '--label' and '--delete-label' values; a runner in group mode keeps a 'group' label", + err, + )) + } agent, err = UpdateAgent(cmd, agent.ID, input) - ui.ExitOnError("updating agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerWriteFailed, + "Error updating the runner", + common2.RunnerWriteHint, + err, + )) + } PrintControlPlaneAgent(*agent) } diff --git a/cmd/kubectl-testkube/commands/agents/utils.go b/cmd/kubectl-testkube/commands/agents/utils.go index 604694ac44a..e56d9be5dfa 100644 --- a/cmd/kubectl-testkube/commands/agents/utils.go +++ b/cmd/kubectl-testkube/commands/agents/utils.go @@ -9,7 +9,6 @@ import ( "fmt" "io" "maps" - "os" "strings" "time" @@ -88,9 +87,9 @@ func (list internalAgents) Table() (header []string, output [][]string) { func (list internalAgents) TableWithEnvironments(showEnvironments bool) (header []string, output [][]string) { if showEnvironments { - header = []string{"Name", "Environment", "Capabilities", "Labels", "Runner Mode", "License", "Agent ID", "Version", "Last Seen"} + header = []string{"Name", "Environment", "Capabilities", "Labels", "Runner Mode", "License", "Runner ID", "Version", "Last Seen"} } else { - header = []string{"Name", "Capabilities", "Labels", "Runner Mode", "License", "Agent ID", "Version", "Last Seen"} + header = []string{"Name", "Capabilities", "Labels", "Runner Mode", "License", "Runner ID", "Version", "Last Seen"} } for _, e := range list { @@ -195,7 +194,7 @@ type unknownAgentsTable struct { // Table implements ui.TableData interface for unknown agents with simplified columns func (t unknownAgentsTable) Table() (header []string, output [][]string) { - header = []string{"Pod Name", "Namespace", "Agent ID", "Org ID", "Env ID", "Ready"} + header = []string{"Pod Name", "Namespace", "Runner ID", "Org ID", "Env ID", "Ready"} for _, e := range t.agents { podName := e.Pod.Name namespace := e.Pod.Namespace @@ -833,15 +832,25 @@ func UiCreateAgent( enableWebhooks bool, ) *cloudclient.Agent { if name == "" { - name = ui.TextInput("agent name") + name = ui.TextInput("runner name") if name == "" { - ui.Failf("agent name is required") + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "No runner name provided", + "Pass the runner name as an argument, for example `testkube create runner my-runner`", + errors.New("runner name is required"), + )) } } // Get existing agent of that name if existing, err := GetControlPlaneAgent(cmd, name); err == nil { - ui.Failf("agent '%s' already exists", existing.Name) + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "Runner already exists", + "Choose a name that is free, or change the existing runner with `testkube update runner`", + errors.Errorf("runner '%s' already exists", existing.Name), + )) } input := cloudclient.AgentInput{ @@ -883,11 +892,25 @@ func UiCreateAgent( } envs, err := GetControlPlaneEnvironments(cmd) - ui.ExitOnError("getting environments", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrEnvResolutionFailed, + "Error getting the environments", + "Check that your credentials are valid and that your user can read the environments of this organization", + err, + )) + } if len(input.Environments) == 0 { cfg, err := config.Load() - ui.ExitOnError("loading config", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrConfigInitFailed, + "Error loading testkube config file", + common2.ConfigFileHint, + err, + )) + } envOpts := []string{envs[cfg.CloudContext.EnvironmentId].Slug} for id := range envs { if id != cfg.CloudContext.EnvironmentId { @@ -913,16 +936,32 @@ func UiCreateAgent( for _, envId := range input.Environments { env, ok := envs[envId] if !ok { - ui.Failf("unknown environment: %s", envId) + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "Unknown environment", + "Pass an environment id or slug of the current organization with '--env', or select one with `testkube set context --env-id `", + errors.Errorf("unknown environment: %s", envId), + )) } if !env.NewArchitecture { - ui.Warn(fmt.Sprintf("Environment '%s' (%s) does not support new architecture. Please upgrade your control plane.", env.Name, env.Id)) - os.Exit(1) + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrInvalidRuntimeParameter, + "Environment does not support the new architecture", + "Upgrade the control plane of this environment, or pass an environment that runs the new architecture with '--env'", + errors.Errorf("environment '%s' (%s) does not support the new architecture", env.Name, env.Id), + )) } } agent, err := CreateAgent(cmd, input) - ui.ExitOnError("creating agent", err) + if err != nil { + common2.HandleCLIError(common2.NewCLIError( + common2.TKErrRunnerWriteFailed, + "Error creating the runner", + common2.RunnerWriteHint, + err, + )) + } PrintControlPlaneAgent(*agent) diff --git a/cmd/kubectl-testkube/commands/aliases_test.go b/cmd/kubectl-testkube/commands/aliases_test.go new file mode 100644 index 00000000000..78b4c07c955 --- /dev/null +++ b/cmd/kubectl-testkube/commands/aliases_test.go @@ -0,0 +1,80 @@ +package commands + +import ( + "testing" + + "github.com/spf13/cobra" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/agent" + "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/agents" + "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/debug" + "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/docker" + "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/pro" +) + +func TestRunnerCommandAliases(t *testing.T) { + t.Parallel() + + tests := []struct { + name string + cmd *cobra.Command + primary string + aliases []string + }{ + {name: "top-level runner", cmd: NewAgentCmd(), primary: "runner", aliases: []string{"agent"}}, + {name: "create runner", cmd: agents.NewCreateAgentCommand(), primary: "runner", aliases: []string{"agent"}}, + {name: "install runner", cmd: agents.NewInstallAgentCommand(), primary: "runner", aliases: []string{"agent"}}, + {name: "get runner", cmd: agents.NewGetAgentCommand(), primary: "runner", aliases: []string{"runners", "agent", "agents", "a"}}, + {name: "delete runner", cmd: agents.NewDeleteAgentCommand(), primary: "runner", aliases: []string{"agent"}}, + {name: "update runner", cmd: agents.NewUpdateAgentCommand(), primary: "runner", aliases: []string{"agent"}}, + {name: "enable runner", cmd: agents.NewEnableAgentCommand(), primary: "runner", aliases: []string{"agent", "gitops"}}, + {name: "disable runner", cmd: agents.NewDisableAgentCommand(), primary: "runner", aliases: []string{"agent", "gitops"}}, + {name: "debug runner", cmd: debug.NewDebugAgentCmd(), primary: "runner", aliases: []string{"agent", "ag", "a"}}, + {name: "migrate runner", cmd: agent.NewMigrateAgentCmd(), primary: "runner", aliases: []string{"agent"}}, + {name: "init runner", cmd: pro.NewInitCmd(), primary: "runner", aliases: []string{"install", "agent", "init"}}, + {name: "docker init", cmd: docker.NewInitCmd(), primary: "init", aliases: []string{"install", "agent", "runner"}}, + } + + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + assertCommandResolves(t, tt.cmd, tt.primary, tt.aliases...) + }) + } +} + +func TestInitProfileAliases(t *testing.T) { + t.Parallel() + + initCmd := NewInitCmd() + assertCommandResolves(t, findSubcommand(t, initCmd, "standalone-runner"), "standalone-runner", "oss", "standalone", "standalone-agent") + assertCommandResolves(t, findSubcommand(t, initCmd, "runner"), "runner", "agent", "install", "init") +} + +func findSubcommand(t *testing.T, parent *cobra.Command, name string) *cobra.Command { + t.Helper() + cmd, _, err := parent.Find([]string{name}) + require.NoError(t, err) + require.NotNil(t, cmd) + return cmd +} + +func assertCommandResolves(t *testing.T, cmd *cobra.Command, primary string, aliases ...string) { + t.Helper() + require.Equal(t, primary, cmd.Name()) + + parent := &cobra.Command{Use: "parent"} + parent.AddCommand(cmd) + + got, _, err := parent.Find([]string{primary}) + require.NoError(t, err) + assert.Equal(t, cmd, got) + + for _, alias := range aliases { + got, _, err = parent.Find([]string{alias}) + require.NoError(t, err, "alias %q should resolve", alias) + assert.Equal(t, cmd, got, "alias %q should resolve to %s", alias, primary) + } +} diff --git a/cmd/kubectl-testkube/commands/common/cloudcontext.go b/cmd/kubectl-testkube/commands/common/cloudcontext.go index 68b80ed1104..2e0a3591d53 100644 --- a/cmd/kubectl-testkube/commands/common/cloudcontext.go +++ b/cmd/kubectl-testkube/commands/common/cloudcontext.go @@ -28,8 +28,8 @@ func UiPrintContext(cfg config.Data) { // add agent information only when need to change agent data, it's usually not needed in usual workflow if ui.IsVerbose() { - contextData["Agent Key"] = text.Obfuscate(cfg.CloudContext.AgentKey) - contextData["Agent URI"] = cfg.CloudContext.AgentUri + contextData["Runner Key"] = text.Obfuscate(cfg.CloudContext.AgentKey) + contextData["Runner URI"] = cfg.CloudContext.AgentUri } ui.InfoGrid(contextData) diff --git a/cmd/kubectl-testkube/commands/common/errors.go b/cmd/kubectl-testkube/commands/common/errors.go index 67be40c9a5b..64ccb029cbf 100644 --- a/cmd/kubectl-testkube/commands/common/errors.go +++ b/cmd/kubectl-testkube/commands/common/errors.go @@ -59,14 +59,18 @@ const ( // TKErrCleanOldMigrationJobFailed is returned in case of issues with old migration jobs. TKErrCleanOldMigrationJobFailed ErrorCode = "TKERR-1401" - // TKERR-15xx errors are related to agent operations. - - // TKErrAgentGetFailed is returned when fetching an agent from the control plane fails. - TKErrAgentGetFailed ErrorCode = "TKERR-1501" - // TKErrAgentRotateKeyFailed is returned when rotating an agent's secret key fails. - TKErrAgentRotateKeyFailed ErrorCode = "TKERR-1502" - // TKErrAgentRotateRegistrationTokenFailed is returned when rotating an environment registration token fails. - TKErrAgentRotateRegistrationTokenFailed ErrorCode = "TKERR-1503" + // TKERR-15xx errors are related to runner operations. + + // TKErrRunnerGetFailed is returned when fetching a runner from the control plane fails. + TKErrRunnerGetFailed ErrorCode = "TKERR-1501" + // TKErrRunnerRotateKeyFailed is returned when rotating a runner's secret key fails. + TKErrRunnerRotateKeyFailed ErrorCode = "TKERR-1502" + // TKErrRunnerRotateRegistrationTokenFailed is returned when rotating an environment registration token fails. + TKErrRunnerRotateRegistrationTokenFailed ErrorCode = "TKERR-1503" + // TKErrRunnerWriteFailed is returned when creating, updating or deleting a runner on the control plane fails. + // Reads use TKErrRunnerGetFailed: the control plane helpers share one preamble and differ only in the final + // call, so a read and a write fail for the same reasons and the code only has to say which was attempted. + TKErrRunnerWriteFailed ErrorCode = "TKERR-1504" // TKERR-16xx errors are related to marketplace operations. @@ -84,6 +88,10 @@ const ( // TKErrResourceNotFound is returned when a requested resource does not exist on the API server. TKErrResourceNotFound ErrorCode = "TKERR-1701" + // TKErrResourceLookupFailed is returned when listing or reading resources from the Kubernetes cluster fails. + // The lookup itself did not complete, so the caller cannot tell whether the resource exists: a resource that + // answered and is absent uses TKErrResourceNotFound. + TKErrResourceLookupFailed ErrorCode = "TKERR-1702" // TKERR-18xx errors are related to authentication and Pro context setup. @@ -99,6 +107,29 @@ const ( TKErrOrgEnvNamesFetchFailed ErrorCode = "TKERR-1805" // TKErrControlPlaneDiscoveryFailed is returned when the Control Plane can't be reached or does not answer with its public info. TKErrControlPlaneDiscoveryFailed ErrorCode = "TKERR-1806" + // TKErrAPIClientInitFailed is returned when the Testkube API client can't be built for a command. + // GetClient fails on the '--header' flag, on the config file, on building the client for the '--client' + // type, and for a cloud context also on refreshing the stored token, on the re-login it falls back to, + // and on writing the new token back. On the default proxy client the kubeconfig is the likely cause and + // the token paths are unreachable, so the hint names the cluster before the credentials. + TKErrAPIClientInitFailed ErrorCode = "TKERR-1807" + + // TKERR-19xx errors are related to the resource commands that talk to the Testkube API. + + // TKErrAPIReadFailed is returned when reading one resource, or listing resources, through the Testkube + // API fails. One code covers get and list for the reason TKErrRunnerWriteFailed gives: the client + // helpers share one preamble and differ only in the final call, so the code says a read was attempted + // and the title says which. A resource the API answered about and that is absent uses + // TKErrResourceNotFound. + TKErrAPIReadFailed ErrorCode = "TKERR-1901" + // TKErrAPIWriteFailed is returned when creating, updating or deleting a resource through the Testkube + // API fails. Reads use TKErrAPIReadFailed. + TKErrAPIWriteFailed ErrorCode = "TKERR-1902" + // TKErrOutputRenderFailed is returned when a command got its result but could not print it: an unusable + // '--output' type or '--go-template' expression, a value that will not marshal, or the CRD template + // behind '--crd-only'. The data is good and only the presentation failed, so the hint points at the + // output flags rather than at the fetch. + TKErrOutputRenderFailed ErrorCode = "TKERR-1903" ) const helpUrl = "https://testkubeworkspace.slack.com" @@ -107,6 +138,67 @@ const helpUrl = "https://testkubeworkspace.slack.com" // the CLI config file. const ConfigFileHint = "Check is the Testkube config file (~/.testkube/config.json) accessible and has right permissions" +// RunnerLookupHint is the recovery hint for a failure to read a runner from the +// control plane. The read fails on the name, on the credentials, or on the +// connection, so the hint names the first two and a command that lists what +// exists. +const RunnerLookupHint = "Check the runner name or ID and that your credentials are valid, or list the runners with `testkube get runners`" + +// RunnerWriteHint is the recovery hint for a failure to create, update or delete +// a runner on the control plane. The preamble is the same as a read, so what is +// left to check is the permission to change it. +const RunnerWriteHint = "Check that your credentials are valid and that your user can manage the runners of this organization" + +// ClusterLookupHint is the recovery hint for a failure to read namespaces, pods +// or CRDs from the cluster. It names the kubeconfig, because the CLI reads the +// cluster with the same context kubectl uses. +const ClusterLookupHint = "Check that your kubeconfig points at the right cluster and that you can read it, for example with `kubectl get namespaces`" + +// APIClientHint is the recovery hint for a failure to build the Testkube API +// client. The client is built from the current context and the stored token, so +// those are the two things to look at. +const APIClientHint = "Check that your kubeconfig points at the right cluster, or sign in again with `testkube pro login` if you use a cloud context and your token has expired" + +// APIReadHint is the recovery hint for a failed read of resources through the +// Testkube API. It does not name a command that lists them: the listing is +// usually the command that just failed. +const APIReadHint = "Check that your credentials are valid and that the current context points at the organization and environment you expect" + +// APIWriteHint is the recovery hint for a failed create, update or delete +// through the Testkube API. A write fails on the same things a read does, so +// what is left to check is the permission to change the resource. +const APIWriteHint = "Check that your credentials are valid and that your user can manage the resources of this environment or namespace" + +// APIDeleteHint is the recovery hint for a failed delete. A delete is the one +// write where the resource being gone already is a normal outcome, which is what +// '--ignore-not-found' is for. +const APIDeleteHint = "Check the name or the '--label' selector, or pass '--ignore-not-found' to succeed when the resource is already gone" + +// OutputRenderHint is the recovery hint for a command that fetched its result +// and could not print it. The fetch worked, so the flags that shape the output +// are what is left. +const OutputRenderHint = "Check the '--output' value (pretty, json, yaml or go) and the '--go-template' expression" + +// NameFlagHint is the recovery hint for a command that needs '--name' and did +// not get one. +const NameFlagHint = "Pass the name with the '--name' flag" + +// NameOrSelectorHint is the recovery hint for a command that accepts either a +// name or a label selector and got neither. +const NameOrSelectorHint = "Pass the name as an argument, or select by labels with '--label', for example '--label app=backend'" + +// NameConflictHint is the recovery hint for a create that found the name taken. +const NameConflictHint = "Choose a name that is free, or pass '--update' to overwrite the existing one" + +// BoolFlagValueHint is the recovery hint for a boolean flag that could not be +// parsed. +const BoolFlagValueHint = "Check the flag value; a boolean flag takes true or false, or drop the flag to use its default" + +// WebhookFlagsHint is the recovery hint for a failure to turn the webhook flags +// into API options. Each of these flags carries structured text the CLI parses, +// which is where such a failure comes from. +const WebhookFlagsHint = "Check the '--events', '--header', '--config' and '--parameter' values, or that '--payload-template' points at a readable file" + type CLIError struct { Code ErrorCode Title string diff --git a/cmd/kubectl-testkube/commands/common/flags.go b/cmd/kubectl-testkube/commands/common/flags.go index a7790477fe6..7224815cf1c 100644 --- a/cmd/kubectl-testkube/commands/common/flags.go +++ b/cmd/kubectl-testkube/commands/common/flags.go @@ -68,7 +68,7 @@ func PopulateMasterFlags(cmd *cobra.Command, opts *HelmOptions, isDockerCmd bool cmd.Flags().BoolVar(&insecure, "cloud-insecure", false, "deprecated: use --skip-tls") cmd.Flags().MarkDeprecated("cloud-insecure", "use --skip-tls") cmd.Flags().StringVar(&agentURIPrefix, "cloud-agent-prefix", defaultAgentPrefix, "usually don't need to be changed [required for custom cloud mode]") - cmd.Flags().MarkDeprecated("cloud-agent-prefix", "use --agent-prefix instead") + cmd.Flags().MarkDeprecated("cloud-agent-prefix", "use --runner-prefix instead") cmd.Flags().StringVar(&apiURIPrefix, "cloud-api-prefix", defaultApiPrefix, "usually don't need to be changed [required for custom cloud mode]") cmd.Flags().MarkDeprecated("cloud-api-prefix", "use --api-prefix instead") cmd.Flags().StringVar(&uiURIPrefix, "cloud-ui-prefix", defaultUiPrefix, "usually don't need to be changed [required for custom cloud mode]") @@ -80,7 +80,9 @@ func PopulateMasterFlags(cmd *cobra.Command, opts *HelmOptions, isDockerCmd bool cmd.Flags().BoolVar(&opts.Master.Insecure, "master-insecure", false, "deprecated: use --skip-tls") cmd.Flags().MarkDeprecated("master-insecure", "use --skip-tls") - cmd.Flags().StringVar(&opts.Master.AgentUrlPrefix, "agent-prefix", defaultAgentPrefix, "usually don't need to be changed [required for custom cloud mode]") + cmd.Flags().StringVar(&opts.Master.AgentUrlPrefix, "runner-prefix", defaultAgentPrefix, "usually don't need to be changed [required for custom cloud mode]") + cmd.Flags().String("agent-prefix", defaultAgentPrefix, "usually don't need to be changed [required for custom cloud mode]") + cmd.Flags().MarkDeprecated("agent-prefix", "use --runner-prefix instead") cmd.Flags().StringVar(&opts.Master.ApiUrlPrefix, "api-prefix", defaultApiPrefix, "usually don't need to be changed [required for custom cloud mode]") cmd.Flags().StringVar(&opts.Master.UiUrlPrefix, "ui-prefix", defaultUiPrefix, "usually don't need to be changed [required for custom cloud mode]") cmd.Flags().StringVar(&opts.Master.RootDomain, "root-domain", defaultRootDomain, "usually don't need to be changed [required for custom cloud mode]") @@ -91,15 +93,21 @@ func PopulateMasterFlags(cmd *cobra.Command, opts *HelmOptions, isDockerCmd bool cmd.Flags().String("api-uri-override", "", "api uri override") cmd.Flags().String("ui-uri-override", "", "ui uri override") cmd.Flags().String("auth-uri-override", "", "auth uri override") + cmd.Flags().String("runner-uri-override", "", "runner uri override") cmd.Flags().String("agent-uri-override", "", "agent uri override") + cmd.Flags().MarkDeprecated("agent-uri-override", "use --runner-uri-override instead") agentURI := "" if isDockerCmd { agentURI = "agent.testkube.io:443" } - cmd.Flags().StringVar(&opts.Master.URIs.Agent, "agent-uri", agentURI, "Testkube Pro agent URI [required for centralized mode]") - cmd.Flags().StringVar(&opts.Master.AgentToken, "agent-token", "", "Testkube Pro agent key [required for centralized mode]") + cmd.Flags().StringVar(&opts.Master.URIs.Agent, "runner-uri", agentURI, "Testkube Pro runner URI [required for centralized mode]") + cmd.Flags().String("agent-uri", agentURI, "Testkube Pro agent URI [required for centralized mode]") + cmd.Flags().MarkDeprecated("agent-uri", "use --runner-uri instead") + cmd.Flags().StringVar(&opts.Master.AgentToken, "runner-token", "", "Testkube Pro runner key [required for centralized mode]") + cmd.Flags().String("agent-token", "", "Testkube Pro agent key [required for centralized mode]") + cmd.Flags().MarkDeprecated("agent-token", "use --runner-token instead") neededForLogin := "" if isDockerCmd { neededForLogin = ". It can be skipped for no login mode" @@ -130,12 +138,18 @@ func ProcessMasterFlags(cmd *cobra.Command, opts *HelmOptions, cfg *config.Data) } } - if !cmd.Flags().Changed("agent-prefix") { - if cmd.Flags().Changed("cloud-agent-prefix") { - opts.Master.AgentUrlPrefix = cmd.Flag("cloud-agent-prefix").Value.String() - } else if configured && cfg.Master.AgentUrlPrefix != "" { - opts.Master.AgentUrlPrefix = cfg.Master.AgentUrlPrefix - } + if v, ok := firstChangedString(cmd, "runner-prefix", "agent-prefix", "cloud-agent-prefix"); ok { + opts.Master.AgentUrlPrefix = v + } else if configured && cfg.Master.AgentUrlPrefix != "" { + opts.Master.AgentUrlPrefix = cfg.Master.AgentUrlPrefix + } + + if v, ok := firstChangedString(cmd, "runner-uri", "agent-uri"); ok { + opts.Master.URIs.Agent = v + } + + if v, ok := firstChangedString(cmd, "runner-token", "agent-token"); ok { + opts.Master.AgentToken = v } if !cmd.Flags().Changed("api-prefix") { @@ -185,8 +199,8 @@ func ProcessMasterFlags(cmd *cobra.Command, opts *HelmOptions, cfg *config.Data) opts.Master.Insecure) // override whole URIs usually composed from prefix - host parts - if flagChanged(cmd, "agent-uri-override") { - uris.WithAgentURI(cmd.Flag("agent-uri-override").Value.String()) + if v, ok := firstChangedString(cmd, "runner-uri-override", "agent-uri-override"); ok { + uris.WithAgentURI(v) } if flagChanged(cmd, "api-uri-override") { @@ -264,6 +278,16 @@ func flagChanged(cmd *cobra.Command, name string) bool { return cmd != nil && cmd.Flag(name) != nil && cmd.Flags().Changed(name) } +// firstChangedString returns the value of the first named flag that was set. +func firstChangedString(cmd *cobra.Command, names ...string) (string, bool) { + for _, name := range names { + if flagChanged(cmd, name) { + return cmd.Flag(name).Value.String(), true + } + } + return "", false +} + // ResolveSkipTLS returns the effective skip-TLS value with precedence: // command flag (--skip-tls, --insecure, --master-insecure, --cloud-insecure) > persisted config > default false. func ResolveSkipTLS(cmd *cobra.Command, cfg *config.Data) bool { @@ -360,22 +384,22 @@ func (s *CommaList) Enabled(value string) bool { func PopulateRunnerFlags(cmd *cobra.Command) { // Installation > General cmd.Flags().StringP("execution-namespace", "N", "", "namespace to run executions (defaults to installation namespace)") - cmd.Flags().String("version", "", "agent version to use (defaults to latest)") + cmd.Flags().String("version", "", "runner version to use (defaults to latest)") cmd.Flags().Bool("dry-run", false, "display helm commands only") // Installation > Runner cmd.Flags().StringP("global-template-path", "g", "", "include global template") - cmd.Flags().Bool("global", false, "make it global agent") - cmd.Flags().String("group", "", "make it grouped agent") + cmd.Flags().Bool("global", false, "make it global runner") + cmd.Flags().String("group", "", "make it grouped runner") // Install existing - cmd.Flags().StringP("secret", "s", "", "secret key for the selected agent") + cmd.Flags().StringP("secret", "s", "", "secret key for the selected runner") // Create and install - cmd.Flags().Bool("create", false, "auto create that agent") - cmd.Flags().StringSliceP("env", "e", nil, "(with --create) environment ID or slug that the agent have access to") + cmd.Flags().Bool("create", false, "auto create that runner") + cmd.Flags().StringSliceP("env", "e", nil, "(with --create) environment ID or slug that the runner have access to") cmd.Flags().StringSliceP("label", "l", nil, "(with --create) label key value pair: --label key1=value1") - cmd.Flags().Bool("floating", false, "(with --create) create as a floating agent") + cmd.Flags().Bool("floating", false, "(with --create) create as a floating runner") // Components selection AddExecutionCapabilityFlags(cmd) @@ -384,7 +408,7 @@ func PopulateRunnerFlags(cmd *cobra.Command) { cmd.Flags().Bool("webhooks", false, "enable webhooks capability") // Deprecated flag - cmd.Flags().StringP("type", "t", "", "[DEPRECATED] agent type - use capability flags instead") + cmd.Flags().StringP("type", "t", "", "[DEPRECATED] runner type - use capability flags instead") cmd.Flags().MarkDeprecated("type", "use --execution, --listener, --gitops, and/or --webhooks instead") } diff --git a/cmd/kubectl-testkube/commands/common/helper.go b/cmd/kubectl-testkube/commands/common/helper.go index 8d768bf4db1..30e2679153b 100644 --- a/cmd/kubectl-testkube/commands/common/helper.go +++ b/cmd/kubectl-testkube/commands/common/helper.go @@ -197,8 +197,8 @@ func decideDemoAgentSecretKey(runnerExists bool, runnerKey string, cpExists bool } return "", false, NewCLIError( TKErrInvalidInstallConfig, - "Existing Testkube demo runner has no readable agent key", - fmt.Sprintf("To fix: recreate the cluster, or delete the %q namespace, then run 'testkube init demo' once. (A demo runner already exists but its agent key can't be read, so this install can't reuse it and a fresh key would not match the Control Plane.)", namespace), + "Existing Testkube demo runner has no readable runner key", + fmt.Sprintf("To fix: recreate the cluster, or delete the %q namespace, then run 'testkube init demo' once. (A demo runner already exists but its runner key can't be read, so this install can't reuse it and a fresh key would not match the Control Plane.)", namespace), fmt.Errorf("demo runner %q found but %s env is empty", demoRunnerDeploymentName, demoRunnerAPIKeyEnvVar), ) } @@ -275,8 +275,8 @@ func HelmUpgradeOrInstallTestkubeAgent(options HelmOptions, cfg config.Data, isM return NewCLIError( TKErrInvalidInstallConfig, "Invalid install config", - "Provide the agent token by setting the '--agent-token' flag", - errors.New("agent key is required")) + "Provide the runner token by setting the '--runner-token' flag", + errors.New("runner key is required")) } if cliErr := updateHelmRepo(helmPath, options.DryRun, false); cliErr != nil { @@ -1459,8 +1459,8 @@ func DockerRunTestkubeAgent(options HelmOptions, cfg config.Data, dockerContaine return NewCLIError( TKErrInvalidInstallConfig, "Invalid install config", - "Provide the agent token by setting the '--agent-token' flag", - errors.New("agent key is required")) + "Provide the runner token by setting the '--runner-token' flag", + errors.New("runner key is required")) } args := prepareTestkubeProDockerArgs(options, dockerContainerName, dockerImage) @@ -1545,7 +1545,7 @@ func StreamDockerLogs(dockerContainerName string) *CLIError { return NewCLIError( TKErrDockerLogStreamingFailed, "Docker log streaming failed", - "Check that your Testkube Docker Agent container is up and runnning", + "Check that your Testkube Docker Runner container is up and runnning", err) } defer logs.Close() @@ -1570,7 +1570,7 @@ func StreamDockerLogs(dockerContainerName string) *CLIError { return NewCLIError( TKErrDockerInstallationFailed, "Docker installation failed", - "Check logs of your Testkube Docker Agent container", + "Check logs of your Testkube Docker Runner container", errors.New(string(line))) } } @@ -1579,7 +1579,7 @@ func StreamDockerLogs(dockerContainerName string) *CLIError { return NewCLIError( TKErrDockerLogReadingFailed, "Docker log reading failed", - "Check logs of your Testkube Docker Agent container", + "Check logs of your Testkube Docker Runner container", err) } @@ -1596,8 +1596,8 @@ func DockerUpgradeTestkubeAgent(options HelmOptions, latestVersion string, cfg c return NewCLIError( TKErrInvalidInstallConfig, "Invalid install config", - "Provide the agent token by setting the '--agent-token' flag", - errors.New("agent key is required")) + "Provide the runner token by setting the '--runner-token' flag", + errors.New("runner key is required")) } args := prepareTestkubeUpgradeDockerArgs(options, cfg.CloudContext.DockerContainerName, latestVersion) diff --git a/cmd/kubectl-testkube/commands/common/masterFlags_test.go b/cmd/kubectl-testkube/commands/common/masterFlags_test.go index 5681109236f..d1177be99a5 100644 --- a/cmd/kubectl-testkube/commands/common/masterFlags_test.go +++ b/cmd/kubectl-testkube/commands/common/masterFlags_test.go @@ -37,14 +37,14 @@ func TestMasterCmds(t *testing.T) { t.Run("Test all master flags set and isnecure", func(t *testing.T) { cmd := NewTestCmd() cmd.SetArgs([]string{"--master-insecure", "true", - "--agent-prefix", "dummy-agent-prefix", + "--runner-prefix", "dummy-agent-prefix", "--api-prefix", "dummy-api-prefix", "--ui-prefix", "dummy-ui-prefix", "--root-domain", "dummy-root-domain", "--dry-run", "true", "--no-confirm", "true", - "--agent-token", "dummy-token", - "--agent-uri", "dummy-uri", + "--runner-token", "dummy-token", + "--runner-uri", "dummy-uri", "--ui-prefix", "dummy-ui-prefix", }) err := cmd.Execute() @@ -63,14 +63,14 @@ func TestMasterCmds(t *testing.T) { }) t.Run("Test all master flags set and secure", func(t *testing.T) { cmd := NewTestCmd() - cmd.SetArgs([]string{"--agent-prefix", "dummy-agent-prefix", + cmd.SetArgs([]string{"--runner-prefix", "dummy-agent-prefix", "--api-prefix", "dummy-api-prefix", "--ui-prefix", "dummy-ui-prefix", "--root-domain", "dummy-root-domain", "--dry-run", "true", "--no-confirm", "true", - "--agent-token", "dummy-token", - "--agent-uri", "dummy-uri", + "--runner-token", "dummy-token", + "--runner-uri", "dummy-uri", "--ui-prefix", "dummy-ui-prefix", }) err := cmd.Execute() @@ -288,9 +288,9 @@ func TestMasterCmds(t *testing.T) { assert.Equal(t, "agent.dummy-root-domain:443", opts.Master.URIs.Agent) }) - t.Run("Test defaults for master flags secure with agent uri modified", func(t *testing.T) { + t.Run("Test defaults for master flags secure with runner uri modified", func(t *testing.T) { cmd := NewTestCmd() - cmd.SetArgs([]string{"--agent-uri", "dummy-agent-uri"}) + cmd.SetArgs([]string{"--runner-uri", "dummy-agent-uri"}) err := cmd.Execute() assert.NoError(t, err) assert.Equal(t, false, opts.Master.Insecure) @@ -306,9 +306,9 @@ func TestMasterCmds(t *testing.T) { assert.Equal(t, "dummy-agent-uri", opts.Master.URIs.Agent) }) - t.Run("Test defaults for master flags insecure with agent uri modified", func(t *testing.T) { + t.Run("Test defaults for master flags insecure with runner uri modified", func(t *testing.T) { cmd := NewTestCmd() - cmd.SetArgs([]string{"--master-insecure", "true", "--agent-uri", "dummy-agent-uri"}) + cmd.SetArgs([]string{"--master-insecure", "true", "--runner-uri", "dummy-agent-uri"}) err := cmd.Execute() assert.NoError(t, err) assert.Equal(t, true, opts.Master.Insecure) @@ -355,6 +355,36 @@ func TestMasterCmds(t *testing.T) { assert.NoError(t, err) assert.Equal(t, "pro-test-domain", opts.Master.RootDomain) }) + + t.Run("deprecated --agent-prefix --agent-token --agent-uri still apply", func(t *testing.T) { + cmd := NewTestCmd() + cmd.SetArgs([]string{ + "--agent-prefix", "legacy-agent-prefix", + "--agent-token", "legacy-token", + "--agent-uri", "legacy-uri", + }) + err := cmd.Execute() + assert.NoError(t, err) + assert.Equal(t, "legacy-token", opts.Master.AgentToken) + assert.Equal(t, "legacy-uri", opts.Master.URIs.Agent) + assert.Equal(t, "legacy-agent-prefix", opts.Master.AgentUrlPrefix) + }) + + t.Run("deprecated --agent-uri-override still applies", func(t *testing.T) { + cmd := NewTestCmd() + cmd.SetArgs([]string{"--agent-uri-override", "https://legacy.example.com:443"}) + err := cmd.Execute() + assert.NoError(t, err) + assert.Equal(t, "https://legacy.example.com:443", opts.Master.URIs.Agent) + }) + + t.Run("--runner-uri-override applies", func(t *testing.T) { + cmd := NewTestCmd() + cmd.SetArgs([]string{"--runner-uri-override", "https://runner.example.com:443"}) + err := cmd.Execute() + assert.NoError(t, err) + assert.Equal(t, "https://runner.example.com:443", opts.Master.URIs.Agent) + }) } func TestMasterOrgEnvNameFlags(t *testing.T) { diff --git a/cmd/kubectl-testkube/commands/context/set.go b/cmd/kubectl-testkube/commands/context/set.go index 425503f552a..639c6a102c7 100644 --- a/cmd/kubectl-testkube/commands/context/set.go +++ b/cmd/kubectl-testkube/commands/context/set.go @@ -223,7 +223,7 @@ func NewSetContextCmd() *cobra.Command { cmd.Flags().StringVarP(&apiKey, "api-key", "k", "", "API Key for Testkube Pro") // allow to override default values of all URIs - cmd.Flags().StringVar(&dockerContainerName, "docker-container", "testkube-agent", "Docker container name for Testkube Docker Agent") + cmd.Flags().StringVar(&dockerContainerName, "docker-container", "testkube-agent", "Docker container name for Testkube Docker Runner") common.PopulateMasterFlags(cmd, &opts, false) diff --git a/cmd/kubectl-testkube/commands/debug/agent.go b/cmd/kubectl-testkube/commands/debug/agent.go index 70195f7b53b..73d6aca831f 100644 --- a/cmd/kubectl-testkube/commands/debug/agent.go +++ b/cmd/kubectl-testkube/commands/debug/agent.go @@ -16,10 +16,10 @@ func NewDebugAgentCmd() *cobra.Command { var show common.CommaList cmd := &cobra.Command{ - Use: "agent", - Aliases: []string{"ag", "a"}, - Short: "Show Agent debug information", - Long: "Get all the necessary information to debug an issue in Testkube Agent you can fiter through comma separated list of items to show with additional flag `--show " + agentFeaturesStr + "`", + Use: "runner", + Aliases: []string{"agent", "ag", "a"}, + Short: "Show Runner debug information", + Long: "Get all the necessary information to debug an issue in Testkube Runner you can fiter through comma separated list of items to show with additional flag `--show " + agentFeaturesStr + "`", Run: RunDebugAgentCmdFunc(&show), } @@ -34,10 +34,10 @@ func RunDebugAgentCmdFunc(show *common.CommaList) func(cmd *cobra.Command, args ui.ExitOnError("loading config file", err) ui.NL() - ui.H1("Agent Insights") + ui.H1("Runner Insights") if cfg.ContextType != config.ContextTypeCloud { - ui.Errf("Agent debug is only available for cloud context") + ui.Errf("Runner debug is only available for cloud context") ui.NL() ui.ShellCommand("Please try command below to set your context into Cloud mode", `testkube set context -o -e -k `) ui.NL() @@ -73,9 +73,9 @@ func RunDebugAgentCmdFunc(show *common.CommaList) func(cmd *cobra.Command, args } if show.Enabled(showApiLogs) { - ui.H2("Agent API Logs") + ui.H2("Runner API Logs") err = common.KubectlLogs(namespace, map[string]string{"app.kubernetes.io/name": "api-server"}) - ui.ExitOnError("getting agent logs", err) + ui.ExitOnError("getting runner logs", err) ui.NL(2) } @@ -95,27 +95,27 @@ func RunDebugAgentCmdFunc(show *common.CommaList) func(cmd *cobra.Command, args ui.ExitOnError("getting client", err) if show.Enabled(showRoundtrip) { - ui.H2("Agent connection through Control Plane from CLI") + ui.H2("Runner connection through Control Plane from CLI") i, err := client.GetServerInfo() if err != nil { - ui.Errf("Error while doing roundtrip to agent: %v", err) + ui.Errf("Error while doing roundtrip to runner: %v", err) ui.NL() ui.Info("Possible reasons:") - ui.Warn("- Please check if your agent organization and environment are set correctly") + ui.Warn("- Please check if your runner organization and environment are set correctly") ui.Warn("- Please check if your API token is set correctly") ui.NL() } else { - ui.Warn("Agent correctly connected to cloud:\n") + ui.Warn("Runner correctly connected to cloud:\n") ui.InfoGrid(map[string]string{ - "Agent version ": i.Version, - "Agent namespace": i.Namespace, + "Runner version ": i.Version, + "Runner namespace": i.Namespace, }) } } if show.Enabled(showCLIToControlPlane) { - ui.H2("Agent connection to Control Plane from CLI") + ui.H2("Runner connection to Control Plane from CLI") debug, err := GetDebugInfo(client) ui.ExitOnError("connecting to Control Plane", err) diff --git a/cmd/kubectl-testkube/commands/debug/oss.go b/cmd/kubectl-testkube/commands/debug/oss.go index e2994bc69f8..e112b4885d9 100644 --- a/cmd/kubectl-testkube/commands/debug/oss.go +++ b/cmd/kubectl-testkube/commands/debug/oss.go @@ -18,7 +18,7 @@ func NewDebugOssCmd() *cobra.Command { ui.ExitOnError("loading config file", err) if cfg.ContextType != config.ContextTypeKubeconfig { - ui.Errf("OSS debug is only available for kubeconfig context, use `testkube set context` to set kubeconfig context, or `testkube debug agent|controlplane` to debug other variants of Testkube") + ui.Errf("OSS debug is only available for kubeconfig context, use `testkube set context` to set kubeconfig context, or `testkube debug runner|controlplane` to debug other variants of Testkube") return } diff --git a/cmd/kubectl-testkube/commands/docker/init.go b/cmd/kubectl-testkube/commands/docker/init.go index 402a9b66712..06f6f7e5327 100644 --- a/cmd/kubectl-testkube/commands/docker/init.go +++ b/cmd/kubectl-testkube/commands/docker/init.go @@ -23,8 +23,8 @@ func NewInitCmd() *cobra.Command { cmd := &cobra.Command{ Use: "init", - Short: "Run Testkube Docker Agent and connect to Testkube Pro environment", - Aliases: []string{"install", "agent"}, + Short: "Run Testkube Docker Runner and connect to Testkube Pro environment", + Aliases: []string{"install", "agent", "runner"}, Run: func(cmd *cobra.Command, args []string) { ui.Info("WELCOME TO") ui.Logo() @@ -60,7 +60,7 @@ func NewInitCmd() *cobra.Command { } if !options.NoConfirm { - ui.Warn("This will run Testkube Docker Agent latest version. This will take a few minutes.") + ui.Warn("This will run Testkube Docker Runner latest version. This will take a few minutes.") ui.Warn("Please be sure you have Docker service running before continuing and can run containers in privileged mode!") ui.NL() @@ -71,7 +71,7 @@ func NewInitCmd() *cobra.Command { ok := ui.Confirm("Do you want to continue?") if !ok { - ui.Errf("Testkube Docker Agent running cancelled") + ui.Errf("Testkube Docker Runner running cancelled") common.SendErrTelemetry(cmd, cfg, "user_cancel", errors.New("user cancelled agent running")) return } @@ -79,9 +79,9 @@ func NewInitCmd() *cobra.Command { var spinner *pterm.SpinnerPrinter if ui.IsVerbose() { - ui.H2("Running Testkube Docker Agent") + ui.H2("Running Testkube Docker Runner") } else { - spinner = ui.NewSpinner("Running Testkube Docker Agent") + spinner = ui.NewSpinner("Running Testkube Docker Runner") } if cliErr := common.DockerRunTestkubeAgent(options, cfg, dockerContainerName, dockerImage); cliErr != nil { @@ -102,7 +102,7 @@ func NewInitCmd() *cobra.Command { if spinner != nil { spinner.Success() } else { - ui.Success("Testkube Docker Agent is up and running") + ui.Success("Testkube Docker Runner is up and running") } if noLogin { @@ -186,8 +186,8 @@ func NewInitCmd() *cobra.Command { common.PopulateMasterFlags(cmd, &options, true) cmd.Flags().BoolVarP(&noLogin, "no-login", "", false, "Ignore login prompt, set existing token later by `testkube set context`") - cmd.Flags().StringVar(&dockerContainerName, "docker-container", "testkube-agent", "Docker container name for Testkube Docker Agent") - cmd.Flags().StringVar(&dockerImage, "docker-image", "kubeshop/testkube-agent:"+StableReleasePlaceholder, "Docker image for Testkube Docker Agent") + cmd.Flags().StringVar(&dockerContainerName, "docker-container", "testkube-agent", "Docker container name for Testkube Docker Runner") + cmd.Flags().StringVar(&dockerImage, "docker-image", "kubeshop/testkube-agent:"+StableReleasePlaceholder, "Docker image for Testkube Docker Runner") return cmd } diff --git a/cmd/kubectl-testkube/commands/init.go b/cmd/kubectl-testkube/commands/init.go index 99dc0d30de1..08854245069 100644 --- a/cmd/kubectl-testkube/commands/init.go +++ b/cmd/kubectl-testkube/commands/init.go @@ -23,14 +23,14 @@ import ( const ( defaultNamespace = "testkube" - standaloneAgentProfile = "standalone-agent" + standaloneAgentProfile = "standalone-runner" demoProfile = "demo" demoValuesUrl = "https://raw.githubusercontent.com/kubeshop/testkube-cloud-charts/main/charts/testkube-enterprise/profiles/values.demo.v2.yaml" - agentProfile = "agent" + agentProfile = "runner" standaloneInstallationName = "Testkube OSS" demoInstallationName = "Testkube On-Prem demo" - agentInstallationName = "Testkube Agent" + agentInstallationName = "Testkube Runner" ) func NewInitCmd() *cobra.Command { @@ -74,7 +74,7 @@ func NewInitCmdStandalone() *cobra.Command { cmd := &cobra.Command{ Use: standaloneAgentProfile, Short: "Install " + standaloneInstallationName + " in your current context", - Aliases: []string{"oss", "standalone"}, + Aliases: []string{"oss", "standalone", "standalone-agent"}, Run: func(cmd *cobra.Command, args []string) { if export { common.HandleCLIError(common.NewCLIError( @@ -265,7 +265,7 @@ func NewInitCmdDemo() *cobra.Command { runnerSecretKey, cliErr := common.ResolveDemoAgentSecretKey(options.Namespace, options.DryRun) if cliErr != nil { spinner.Fail("Failed to install Testkube On-Prem Demo") - exitOnInstallError(cmd, cfg, "resolving agent key", "install_failed", license, cliErr) + exitOnInstallError(cmd, cfg, "resolving runner key", "install_failed", license, cliErr) } cliErr = common.HelmUpgradeOrInstallTestkubeOnPremDemo(options, runnerSecretKey) diff --git a/cmd/kubectl-testkube/commands/pro/connect.go b/cmd/kubectl-testkube/commands/pro/connect.go index ddac4d1c97c..cc08dd14153 100644 --- a/cmd/kubectl-testkube/commands/pro/connect.go +++ b/cmd/kubectl-testkube/commands/pro/connect.go @@ -36,7 +36,7 @@ func NewConnectCmd() *cobra.Command { ) cmd := &cobra.Command{ - Use: "connect [agent-name]", + Use: "connect [runner-name]", Aliases: []string{"c"}, Args: cobra.MaximumNArgs(1), Short: "Testkube Pro connect ", @@ -180,10 +180,10 @@ func NewConnectCmd() *cobra.Command { ui.Failf("You need pass valid organization id to connect to Pro") } if masterOpts.Master.URIs.Agent == "" { - ui.Failf("You need pass valid uri agent to connect to Pro") + ui.Failf("You need pass valid runner uri to connect to Pro") } - // Export execution data before switching to agent mode + // Export execution data before switching to runner mode var exportPath string var exportDir string if !skipExport { @@ -237,7 +237,7 @@ func NewConnectCmd() *cobra.Command { err = config.Save(cfg) ui.ExitOnError("saving cloud context configuration", err) - // Install agent using same mechanism as "install agent" command + // Install runner using same mechanism as "install runner" command agentName := "default-oss" if len(args) > 0 && args[0] != "" { agentName = args[0] @@ -253,7 +253,7 @@ func NewConnectCmd() *cobra.Command { _ = cmd.Flags().Set("execution", "true") } - ui.H2("Switching OSS Standalone Agent to Cloud Runner mode") + ui.H2("Switching OSS Standalone Runner to Cloud mode") agents.UiInstallAgent(cmd, agentName, []string{"testkube.io/source=oss"}, map[string]interface{}{ // Disable CRD installation in the runner chart — the OSS chart already @@ -370,7 +370,7 @@ func NewConnectCmd() *cobra.Command { cmd.Flags().BoolVar(&skipExport, "skip-export", false, "Skip exporting execution data before connecting") cmd.Flags().StringVar(&exportSince, "since", "", "Export only executions created after this date (e.g. 2025-01-01 or 2025-01-01T00:00:00Z)") - // Cloud/master flags (--org-id, --env-id, --root-domain, --agent-token, etc.) + // Cloud/master flags (--org-id, --env-id, --root-domain, --runner-token, etc.) cmd.Flags().StringVarP(&apiKey, "api-key", "k", "", "API Key for Testkube Pro") common.PopulateMasterFlags(cmd, &masterOpts, false) diff --git a/cmd/kubectl-testkube/commands/pro/disconnect.go b/cmd/kubectl-testkube/commands/pro/disconnect.go index edad60e5133..bf0d13cb9ca 100644 --- a/cmd/kubectl-testkube/commands/pro/disconnect.go +++ b/cmd/kubectl-testkube/commands/pro/disconnect.go @@ -71,7 +71,7 @@ func NewDisconnectCmd() *cobra.Command { // uninstall the runner chart that was installed by "pro connect"; // failures are non-fatal so disconnect can still restore OSS mode if cfg.CloudContext.AgentReleaseName != "" && cfg.CloudContext.AgentNamespace != "" { - spinner := ui.NewSpinner("Uninstalling agent runner") + spinner := ui.NewSpinner("Uninstalling runner") if cliErr := common.HelmUninstall(cfg.CloudContext.AgentNamespace, cfg.CloudContext.AgentReleaseName); cliErr != nil { spinner.Fail(fmt.Sprintf("Failed to uninstall runner release %s (continuing with disconnect): %s", cfg.CloudContext.AgentReleaseName, cliErr)) } else { @@ -81,9 +81,9 @@ func NewDisconnectCmd() *cobra.Command { // Delete the agent record from the control plane that was created by "pro connect" if cfg.CloudContext.AgentName != "" && cfg.CloudContext.ApiUri != "" && cfg.CloudContext.ApiKey != "" && cfg.CloudContext.OrganizationId != "" { - spinner := ui.NewSpinner("Deleting agent from control plane") + spinner := ui.NewSpinner("Deleting runner from control plane") if err := common.DeleteAgent(cfg.CloudContext.ApiUri, cfg.CloudContext.ApiKey, cfg.CloudContext.OrganizationId, cfg.CloudContext.AgentName, skipTLS); err != nil { - spinner.Fail(fmt.Sprintf("Failed to delete agent %q from control plane (continuing with disconnect): %s", cfg.CloudContext.AgentName, err)) + spinner.Fail(fmt.Sprintf("Failed to delete runner %q from control plane (continuing with disconnect): %s", cfg.CloudContext.AgentName, err)) } else { spinner.Success() } diff --git a/cmd/kubectl-testkube/commands/pro/init.go b/cmd/kubectl-testkube/commands/pro/init.go index e9c99f247dc..311e04d8ef9 100644 --- a/cmd/kubectl-testkube/commands/pro/init.go +++ b/cmd/kubectl-testkube/commands/pro/init.go @@ -22,15 +22,15 @@ func NewInitCmd() *cobra.Command { } cmd := &cobra.Command{ - Use: "agent", - Short: "Install Testkube Pro Agent and connect to Testkube Pro environment", + Use: "runner", + Short: "Install Testkube Pro Runner and connect to Testkube Pro environment", Aliases: []string{"install", "agent", "init"}, Run: func(cmd *cobra.Command, args []string) { if export { common.HandleCLIError(common.NewCLIError( common.TKErrInvalidRuntimeParameter, "Export is unavailable for this profile", - "Drop the '--export' flag, it is only supported when installing the standalone agent", + "Drop the '--export' flag, it is only supported when installing the standalone runner", errors.New("export is unavailable for this profile"), )) } @@ -52,7 +52,7 @@ func NewInitCmd() *cobra.Command { skipTLS := common.SyncSkipTLSFromFlags(cmd, &cfg) common.ProcessMasterFlags(cmd, &options, &cfg) - common.ShowOperatorDeprecationWarning("Testkube Agent", options.NoCRDs) + common.ShowOperatorDeprecationWarning("Testkube Runner", options.NoCRDs) common.SendAttemptTelemetry(cmd, cfg) diff --git a/cmd/kubectl-testkube/commands/pro/login.go b/cmd/kubectl-testkube/commands/pro/login.go index 9a526f813f6..de2ccac1423 100644 --- a/cmd/kubectl-testkube/commands/pro/login.go +++ b/cmd/kubectl-testkube/commands/pro/login.go @@ -140,11 +140,11 @@ func NewLoginCmd() *cobra.Command { if !cmd.Flags().Changed("ui-uri-override") && result.UIURL != "" { cmd.Flags().Set("ui-uri-override", result.UIURL) } - if !cmd.Flags().Changed("agent-uri-override") && result.AgentURL != "" { + if !cmd.Flags().Changed("runner-uri-override") && !cmd.Flags().Changed("agent-uri-override") && result.AgentURL != "" { if !strings.Contains(result.AgentURL, "://") { result.AgentURL = fmt.Sprintf("%s://%s", u.Scheme, result.AgentURL) } - cmd.Flags().Set("agent-uri-override", result.AgentURL) + cmd.Flags().Set("runner-uri-override", result.AgentURL) } if !cmd.Flags().Changed("callback-port") { diff --git a/cmd/kubectl-testkube/commands/status.go b/cmd/kubectl-testkube/commands/status.go index 21acf2fd067..3d12e12de80 100644 --- a/cmd/kubectl-testkube/commands/status.go +++ b/cmd/kubectl-testkube/commands/status.go @@ -74,7 +74,7 @@ func printContextStatus(cmd *cobra.Command, cfg config.Data) { {"Namespace ", cfg.Namespace}, }) } else { - ui.PrintEnabled("Context", "connected to a local standalone agent") + ui.PrintEnabled("Context", "connected to a local standalone runner") namespace := cfg.Namespace if flag := cmd.Flag("namespace"); flag != nil && flag.Changed { diff --git a/cmd/kubectl-testkube/commands/testworkflows/get.go b/cmd/kubectl-testkube/commands/testworkflows/get.go index 4c75af34c9f..5e378cc0fbc 100644 --- a/cmd/kubectl-testkube/commands/testworkflows/get.go +++ b/cmd/kubectl-testkube/commands/testworkflows/get.go @@ -32,7 +32,7 @@ func NewGetTestWorkflowsCmd() *cobra.Command { Aliases: []string{"testworkflows", "tw"}, Args: cobra.MaximumNArgs(1), Short: "Get all available test workflows", - Long: `Get all available test workflows. In cloud context (API key) the CLI fetches them from the connected Control Plane environment and ignores the namespace flag. In kubeconfig context it fetches them from the agent in the given namespace (default "testkube").`, + Long: `Get all available test workflows. In cloud context (API key) the CLI fetches them from the connected Control Plane environment and ignores the namespace flag. In kubeconfig context it fetches them from the runner in the given namespace (default "testkube").`, PreRunE: func(cmd *cobra.Command, args []string) error { if limit < 0 { diff --git a/cmd/kubectl-testkube/commands/testworkflowtemplates/get.go b/cmd/kubectl-testkube/commands/testworkflowtemplates/get.go index e355dd9ea92..43097ce1394 100644 --- a/cmd/kubectl-testkube/commands/testworkflowtemplates/get.go +++ b/cmd/kubectl-testkube/commands/testworkflowtemplates/get.go @@ -27,7 +27,7 @@ func NewGetTestWorkflowTemplatesCmd() *cobra.Command { Aliases: []string{"testworkflowtemplates", "twt"}, Args: cobra.MaximumNArgs(1), Short: "Get all available test workflow templates", - Long: `Get all available test workflow templates. In cloud context (API key) the CLI fetches them from the connected Control Plane environment and ignores the namespace flag. In kubeconfig context it fetches them from the agent in the given namespace (default "testkube").`, + Long: `Get all available test workflow templates. In cloud context (API key) the CLI fetches them from the connected Control Plane environment and ignores the namespace flag. In kubeconfig context it fetches them from the runner in the given namespace (default "testkube").`, Run: func(cmd *cobra.Command, args []string) { namespace := cmd.Flag("namespace").Value.String() diff --git a/cmd/kubectl-testkube/commands/upgrade.go b/cmd/kubectl-testkube/commands/upgrade.go index 86e5d752da9..fc84a4227d2 100644 --- a/cmd/kubectl-testkube/commands/upgrade.go +++ b/cmd/kubectl-testkube/commands/upgrade.go @@ -67,8 +67,8 @@ func NewUpgradeCmd() *cobra.Command { if ui.IsVerbose() && cfg.ContextType == config.ContextTypeCloud { ui.Info("Your Testkube is in 'cloud' mode with following context") ui.InfoGrid(map[string]string{ - "Agent Key": text.Obfuscate(cfg.CloudContext.AgentKey), - "Agent URI": cfg.CloudContext.AgentUri, + "Runner Key": text.Obfuscate(cfg.CloudContext.AgentKey), + "Runner URI": cfg.CloudContext.AgentUri, }) ui.NL() } @@ -81,7 +81,7 @@ func NewUpgradeCmd() *cobra.Command { } if cfg.ContextType == config.ContextTypeCloud { - ui.Info("Testkube Pro agent upgrade started") + ui.Info("Testkube Pro runner upgrade started") // Both upgrade helpers return *CLIError, so the result has to stay // typed: assigning a nil *CLIError to an error turns it into a // non-nil interface and reports a successful upgrade as a failure. @@ -105,7 +105,7 @@ func NewUpgradeCmd() *cobra.Command { if err = common.PopulateAgentDataToContext(options, cfg); err != nil { common.HandleCLIError(common.NewCLIError( common.TKErrContextSaveFailed, - "Error storing agent data in context", + "Error storing runner data in context", common.ConfigFileHint, err, )) @@ -122,7 +122,7 @@ func NewUpgradeCmd() *cobra.Command { common.PopulateHelmFlags(cmd, &options) common.PopulateMasterFlags(cmd, &options, false) - cmd.Flags().StringVar(&dockerContainerName, "docker-container", "testkube-agent", "Docker container name for Testkube Docker Agent") + cmd.Flags().StringVar(&dockerContainerName, "docker-container", "testkube-agent", "Docker container name for Testkube Docker Runner") return cmd } diff --git a/cmd/kubectl-testkube/commands/webhooks/common.go b/cmd/kubectl-testkube/commands/webhooks/common.go index 36a9955dc1a..b1a6e51e68c 100644 --- a/cmd/kubectl-testkube/commands/webhooks/common.go +++ b/cmd/kubectl-testkube/commands/webhooks/common.go @@ -11,7 +11,6 @@ import ( apiv1 "github.com/kubeshop/testkube/pkg/api/v1/client" "github.com/kubeshop/testkube/pkg/api/v1/testkube" webhooksmapper "github.com/kubeshop/testkube/pkg/mapper/webhooks" - "github.com/kubeshop/testkube/pkg/ui" ) // NewCreateWebhookOptionsFromFlags creates create webhook options from command flags @@ -28,7 +27,9 @@ func NewCreateWebhookOptionsFromFlags(cmd *cobra.Command) (options apiv1.CreateW payloadTemplateContent := "" if payloadTemplate != "" { b, err := os.ReadFile(payloadTemplate) - ui.ExitOnError("reading payload template", err) + if err != nil { + return options, fmt.Errorf("reading payload template: %w", err) + } payloadTemplateContent = string(b) } @@ -142,7 +143,7 @@ func NewUpdateWebhookOptionsFromFlags(cmd *cobra.Command) (options apiv1.UpdateW payloadTemplate := cmd.Flag("payload-template").Value.String() b, err := os.ReadFile(payloadTemplate) if err != nil { - return options, fmt.Errorf("reading payload template %w", err) + return options, fmt.Errorf("reading payload template: %w", err) } value := string(b) diff --git a/cmd/kubectl-testkube/commands/webhooks/create.go b/cmd/kubectl-testkube/commands/webhooks/create.go index 5f5d1c10b27..64a8ecf0d45 100644 --- a/cmd/kubectl-testkube/commands/webhooks/create.go +++ b/cmd/kubectl-testkube/commands/webhooks/create.go @@ -1,6 +1,7 @@ package webhooks import ( + "errors" "fmt" "strconv" @@ -37,23 +38,61 @@ func NewCreateWebhookCmd() *cobra.Command { Long: `Create new Webhook Custom Resource`, Run: func(cmd *cobra.Command, args []string) { crdOnly, err := strconv.ParseBool(cmd.Flag("crd-only").Value.String()) - ui.ExitOnError("parsing flag value", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "Error reading the crd-only flag", + common.BoolFlagValueHint, + err, + )) + } if name == "" { - ui.Failf("pass valid name (in '--name' flag)") + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "No webhook name provided", + common.NameFlagHint, + errors.New("no webhook name provided"), + )) } - namespace := cmd.Flag("namespace").Value.String() - var client apiv1.Client + // The namespace comes back from GetClient; the flag is read there, not here. + var ( + client apiv1.Client + namespace string + ) if !crdOnly { client, namespace, err = common.GetClient(cmd) - ui.ExitOnError("getting client", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIClientInitFailed, + "Error creating the Testkube API client", + common.APIClientHint, + err, + )) + } - webhook, _ := client.GetWebhook(name) - if name == webhook.Name { + // A 404 means there is nothing to overwrite and creation carries on below. Any other + // failure means the lookup never answered, so stop instead of silently creating. + webhook, err := client.GetWebhook(name) + if err != nil && !apiv1.IsNotFound(err) { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIReadFailed, + "Error checking whether the webhook already exists", + common.APIReadHint, + err, + )) + } + + if err == nil && name == webhook.Name { if cmd.Flag("update").Changed { if !update { - ui.Failf("Webhook with name '%s' already exists in namespace %s, ", webhook.Name, namespace) + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "Webhook already exists", + common.NameConflictHint, + fmt.Errorf("webhook '%s' already exists in namespace '%s'", webhook.Name, namespace), + )) } } else { ok := ui.Confirm(fmt.Sprintf("Webhook with name '%s' already exists in namespace %s, ", webhook.Name, namespace) + @@ -64,28 +103,63 @@ func NewCreateWebhookCmd() *cobra.Command { } options, err := NewUpdateWebhookOptionsFromFlags(cmd) - ui.ExitOnError("getting webhook options", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "Error reading the webhook flags", + common.WebhookFlagsHint, + err, + )) + } _, err = client.UpdateWebhook(options) - ui.ExitOnError("updating webhook "+name+" in namespace "+namespace, err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIWriteFailed, + "Error updating the webhook", + common.APIWriteHint, + err, + )) + } ui.SuccessAndExit("Webhook updated", name) } } options, err := NewCreateWebhookOptionsFromFlags(cmd) - ui.ExitOnError("getting webhook options", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "Error reading the webhook flags", + common.WebhookFlagsHint, + err, + )) + } if !crdOnly { _, err := client.CreateWebhook(options) - ui.ExitOnError("creating webhook "+name+" in namespace "+namespace, err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIWriteFailed, + "Error creating the webhook", + common.APIWriteHint, + err, + )) + } ui.Success("Webhook created", name) } else { (*testkube.WebhookCreateRequest)(&options).QuoteTextFields() data, err := crd.ExecuteTemplate(crd.TemplateWebhook, options) - ui.ExitOnError("executing crd template", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrOutputRenderFailed, + "Error rendering the webhook CRD", + "Check the flag values that go into the CRD, or drop '--crd-only' to create the webhook through the Testkube API", + err, + )) + } fmt.Print(data) } diff --git a/cmd/kubectl-testkube/commands/webhooks/delete.go b/cmd/kubectl-testkube/commands/webhooks/delete.go index 62b6f6bbf94..e6c2062eb8f 100644 --- a/cmd/kubectl-testkube/commands/webhooks/delete.go +++ b/cmd/kubectl-testkube/commands/webhooks/delete.go @@ -1,6 +1,7 @@ package webhooks import ( + "errors" "strings" "github.com/spf13/cobra" @@ -21,10 +22,24 @@ func NewDeleteWebhookCmd() *cobra.Command { Long: `Delete webhook, pass webhook name which should be deleted`, Run: func(cmd *cobra.Command, args []string) { ignoreNotFound, err := cmd.Flags().GetBool("ignore-not-found") - ui.ExitOnError("reading flag ignore-not-found", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "Error reading the ignore-not-found flag", + common.BoolFlagValueHint, + err, + )) + } client, _, err := common.GetClient(cmd) - ui.ExitOnError("getting client", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIClientInitFailed, + "Error creating the Testkube API client", + common.APIClientHint, + err, + )) + } if len(args) > 0 { name = args[0] @@ -33,7 +48,14 @@ func NewDeleteWebhookCmd() *cobra.Command { ui.Info("Webhook '" + name + "' not found, but ignoring since --ignore-not-found was passed") ui.SuccessAndExit("Operation completed") } - ui.ExitOnError("deleting webhook: "+name, err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIWriteFailed, + "Error deleting the webhook", + common.APIDeleteHint, + err, + )) + } ui.SuccessAndExit("Successfully deleted webhook", name) } @@ -44,11 +66,23 @@ func NewDeleteWebhookCmd() *cobra.Command { ui.Info("Webhook not found for matching selector '" + selector + "', but ignoring since --ignore-not-found was passed") ui.SuccessAndExit("Operation completed") } - ui.ExitOnError("deleting webhooks by labels: "+selector, err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIWriteFailed, + "Error deleting the webhooks", + common.APIDeleteHint, + err, + )) + } ui.SuccessAndExit("Successfully deleted webhooks by labels", selector) } - ui.Failf("Pass Webhook name or labels to delete by labels") + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "No webhook name or label selector provided", + common.NameOrSelectorHint, + errors.New("no webhook name or label selector provided"), + )) }, } diff --git a/cmd/kubectl-testkube/commands/webhooks/get.go b/cmd/kubectl-testkube/commands/webhooks/get.go index 67d52a5c15a..71575d7d8d8 100644 --- a/cmd/kubectl-testkube/commands/webhooks/get.go +++ b/cmd/kubectl-testkube/commands/webhooks/get.go @@ -8,8 +8,8 @@ import ( "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/common" "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/common/render" + apiclient "github.com/kubeshop/testkube/pkg/api/v1/client" "github.com/kubeshop/testkube/pkg/crd" - "github.com/kubeshop/testkube/pkg/ui" ) func NewGetWebhookCmd() *cobra.Command { @@ -24,13 +24,37 @@ func NewGetWebhookCmd() *cobra.Command { Long: `Get webhook, you can change output format, to get single details pass name as first arg`, Run: func(cmd *cobra.Command, args []string) { client, _, err := common.GetClient(cmd) - ui.ExitOnError("getting client", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIClientInitFailed, + "Error creating the Testkube API client", + common.APIClientHint, + err, + )) + } firstEntry := true if len(args) > 0 { name := args[0] webhook, err := client.GetWebhook(name) - ui.ExitOnError("getting webhook: "+name, err) + if err != nil { + // The API answered that the webhook is absent, which is a different thing to tell the + // user than a read that did not complete. + if apiclient.IsNotFound(err) { + common.HandleCLIError(common.NewCLIError( + common.TKErrResourceNotFound, + "Webhook not found", + "Check the webhook name, or list the webhooks with `testkube get webhooks`", + err, + )) + } + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIReadFailed, + "Error getting the webhook", + common.APIReadHint, + err, + )) + } if crdOnly { webhook.QuoteTextFields() @@ -39,10 +63,24 @@ func NewGetWebhookCmd() *cobra.Command { } err = render.Obj(cmd, webhook, os.Stdout) - ui.ExitOnError("rendering obj", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrOutputRenderFailed, + "Error rendering the webhook", + common.OutputRenderHint, + err, + )) + } } else { webhooks, err := client.ListWebhooks(strings.Join(selectors, ",")) - ui.ExitOnError("getting webhooks", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIReadFailed, + "Error getting the webhooks", + common.APIReadHint, + err, + )) + } if crdOnly { for _, webhook := range webhooks { @@ -54,7 +92,14 @@ func NewGetWebhookCmd() *cobra.Command { } err = render.List(cmd, webhooks, os.Stdout) - ui.ExitOnError("rendering list", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrOutputRenderFailed, + "Error rendering the webhooks", + common.OutputRenderHint, + err, + )) + } } }, } diff --git a/cmd/kubectl-testkube/commands/webhooks/update.go b/cmd/kubectl-testkube/commands/webhooks/update.go index c5da4868714..e308e983e3a 100644 --- a/cmd/kubectl-testkube/commands/webhooks/update.go +++ b/cmd/kubectl-testkube/commands/webhooks/update.go @@ -1,9 +1,13 @@ package webhooks import ( + "errors" + "fmt" + "github.com/spf13/cobra" "github.com/kubeshop/testkube/cmd/kubectl-testkube/commands/common" + apiclient "github.com/kubeshop/testkube/pkg/api/v1/client" "github.com/kubeshop/testkube/pkg/ui" ) @@ -30,22 +34,63 @@ func UpdateWebhookCmd() *cobra.Command { Long: `Update Webhook Custom Resource`, Run: func(cmd *cobra.Command, args []string) { if name == "" { - ui.Failf("pass valid name (in '--name' flag)") + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "No webhook name provided", + common.NameFlagHint, + errors.New("no webhook name provided"), + )) } client, namespace, err := common.GetClient(cmd) - ui.ExitOnError("getting client", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIClientInitFailed, + "Error creating the Testkube API client", + common.APIClientHint, + err, + )) + } - webhook, _ := client.GetWebhook(name) - if name != webhook.Name { - ui.Failf("Webhook with name '%s' not exists in namespace %s", name, namespace) + _, err = client.GetWebhook(name) + if err != nil { + // The API answered that the webhook is absent, which is a different thing to tell the + // user than a read that did not complete. + if apiclient.IsNotFound(err) { + common.HandleCLIError(common.NewCLIError( + common.TKErrResourceNotFound, + "Webhook not found", + "Check the '--name' and '--namespace' values, or create the webhook with `testkube create webhook`", + fmt.Errorf("webhook '%s' not found in namespace '%s': %w", name, namespace, err), + )) + } + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIReadFailed, + "Error getting the webhook", + common.APIReadHint, + err, + )) } options, err := NewUpdateWebhookOptionsFromFlags(cmd) - ui.ExitOnError("getting webhook options", err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrInvalidRuntimeParameter, + "Error reading the webhook flags", + common.WebhookFlagsHint, + err, + )) + } _, err = client.UpdateWebhook(options) - ui.ExitOnError("updating webhook "+name+" in namespace "+namespace, err) + if err != nil { + common.HandleCLIError(common.NewCLIError( + common.TKErrAPIWriteFailed, + "Error updating the webhook", + common.APIWriteHint, + err, + )) + } ui.Success("Webhook updated", name) }, diff --git a/go.mod b/go.mod index a8cb34320e1..b72457b26c7 100644 --- a/go.mod +++ b/go.mod @@ -43,7 +43,7 @@ require ( github.com/kelseyhightower/envconfig v1.4.0 github.com/keygen-sh/jsonapi-go v1.2.1 github.com/keygen-sh/keygen-go/v3 v3.3.0 - github.com/klauspost/compress v1.20.0 + github.com/klauspost/compress v1.20.1 github.com/mark3labs/mcp-go v1.1.1 github.com/mattn/go-isatty v0.0.24 github.com/minio/minio-go/v7 v7.3.0 @@ -52,7 +52,7 @@ require ( github.com/ohler55/ojg v1.28.6 github.com/olekukonko/tablewriter v1.1.5 github.com/onsi/ginkgo/v2 v2.33.0 - github.com/onsi/gomega v1.43.1 + github.com/onsi/gomega v1.44.0 github.com/otiai10/copy v1.14.1 github.com/pashagolub/pgxmock/v5 v5.2.0 github.com/pkg/errors v0.9.1 diff --git a/go.sum b/go.sum index 0a4ed808b07..b0ea2ded060 100644 --- a/go.sum +++ b/go.sum @@ -404,6 +404,8 @@ github.com/keygen-sh/keygen-go/v3 v3.3.0 h1:EuX4HIxGgkfnF6GKOolvTJhch1SRpKq7bIxp github.com/keygen-sh/keygen-go/v3 v3.3.0/go.mod h1:YoFyryzXEk6XrbT3H8EUUU+JcIJkQu414TA6CvZgS/E= github.com/klauspost/compress v1.20.0 h1:a3C1ke2ohxFymNlb2HWAHjDeKCI90scRskErZkR0ezA= github.com/klauspost/compress v1.20.0/go.mod h1:LUdAzn7YLVvxLpc7y3V1m40wESHTgc1422pwwBSKYuI= +github.com/klauspost/compress v1.20.1 h1:T7kKElXUMXrUJ2E9QhQhxFtcK5rPyLdsGZvdbLMPdiQ= +github.com/klauspost/compress v1.20.1/go.mod h1:LUdAzn7YLVvxLpc7y3V1m40wESHTgc1422pwwBSKYuI= github.com/klauspost/cpuid/v2 v2.0.1/go.mod h1:FInQzS24/EEf25PyTYn52gqo7WaD8xa0213Md/qVLRg= github.com/klauspost/cpuid/v2 v2.0.9/go.mod h1:FInQzS24/EEf25PyTYn52gqo7WaD8xa0213Md/qVLRg= github.com/klauspost/cpuid/v2 v2.0.10/go.mod h1:g2LTdtYhdyuGPqyWyv7qRAmj1WBqxuObKfj5c0PQa7c= @@ -538,6 +540,8 @@ github.com/onsi/gomega v1.43.0 h1:VlG/1FxqNxhSO+lq/OHBNaaqwiBK/mO8JbVkX9Y+FeU= github.com/onsi/gomega v1.43.0/go.mod h1:REff/hsDsodHoKlWsP2mAPhu1+5/6hVYNf9rIEBpeSg= github.com/onsi/gomega v1.43.1 h1:vGIPFuYrIO6/0Z09s0I0QQQgFchiX4+tb1re3MScJYo= github.com/onsi/gomega v1.43.1/go.mod h1:e/C2HwaZ1DhvjzXXuFhcR7hY7Sh9pl7MmoWKEjzwcdA= +github.com/onsi/gomega v1.44.0 h1:eAiGl3Pw5jz5GQdDff0BcxYpAX1JxW8xD7mFUuwNfZQ= +github.com/onsi/gomega v1.44.0/go.mod h1:e/C2HwaZ1DhvjzXXuFhcR7hY7Sh9pl7MmoWKEjzwcdA= github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U= github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM= github.com/opencontainers/image-spec v1.1.1 h1:y0fUlFfIZhPF1W537XOLg0/fcx6zcHCJwooC2xJA040= diff --git a/internal/config/procontext.go b/internal/config/procontext.go index fbf59ee834a..fa085f8b86b 100644 --- a/internal/config/procontext.go +++ b/internal/config/procontext.go @@ -90,10 +90,17 @@ func (a *ProContextAgent) HasCapability(capability cloud.AgentCapability) bool { } // ShouldPushClusterInventory reports whether this agent should run the CRD -// watcher and push the cluster-resources inventory. Only listener-capable -// agents should: the Control Plane rejects everyone else's push, and a -// runner-only deployment has no CRD RBAC to watch with. Standalone serves -// discovery from its own API, so it never pushes. -func ShouldPushClusterInventory(proContext ProContext) bool { - return proContext.APIKey != "" && proContext.Agent.HasCapability(cloud.AgentCapability_AGENT_CAPABILITY_LISTENER) +// watcher and push the cluster-resources inventory. See internal/inventory. +func ShouldPushClusterInventory(proContext ProContext, testTriggersDisabled bool) bool { + // The inventory only feeds the TestTrigger resourceRef picker. + if testTriggersDisabled { + return false + } + // Standalone serves discovery from its own API, so it never pushes. + if proContext.APIKey == "" { + return false + } + // The Control Plane rejects a push from anyone but a listener, and a + // runner-only deployment has no CRD RBAC to watch with. + return proContext.Agent.HasCapability(cloud.AgentCapability_AGENT_CAPABILITY_LISTENER) } diff --git a/internal/config/procontext_test.go b/internal/config/procontext_test.go index cc6bda305ff..fba24336f06 100644 --- a/internal/config/procontext_test.go +++ b/internal/config/procontext_test.go @@ -19,9 +19,10 @@ func TestShouldPushClusterInventory(t *testing.T) { cloud.AgentCapability_AGENT_CAPABILITY_EXECUTION, } tests := []struct { - name string - proContext ProContext - want bool + name string + proContext ProContext + testTriggersDisabled bool + want bool }{ { name: "standalone mode never pushes", @@ -33,6 +34,12 @@ func TestShouldPushClusterInventory(t *testing.T) { proContext: ProContext{APIKey: "key", Agent: ProContextAgent{Capabilities: listener}}, want: true, }, + { + name: "connected listener-capable agent with test triggers disabled does not push", + proContext: ProContext{APIKey: "key", Agent: ProContextAgent{Capabilities: listener}}, + testTriggersDisabled: true, + want: false, + }, { name: "connected execution-only agent does not push", proContext: ProContext{APIKey: "key", Agent: ProContextAgent{Capabilities: executionOnly}}, @@ -46,7 +53,7 @@ func TestShouldPushClusterInventory(t *testing.T) { } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { - assert.Equal(t, tt.want, ShouldPushClusterInventory(tt.proContext)) + assert.Equal(t, tt.want, ShouldPushClusterInventory(tt.proContext, tt.testTriggersDisabled)) }) } } diff --git a/internal/inventory/doc.go b/internal/inventory/doc.go new file mode 100644 index 00000000000..d5ac73f2d53 --- /dev/null +++ b/internal/inventory/doc.go @@ -0,0 +1,26 @@ +// Package inventory is the Agent side of AgentInventoryService: it publishes a +// snapshot of the resource kinds the Agent can watch in its cluster to the +// Control Plane, which caches it to render the TestTrigger resourceRef +// picker. It is +// cluster-environment metadata, not Testkube objects - those travel through +// internal/sync. +// +// - controller runs the push loop (startup, an hourly safety net, and +// debounced events from a CRD informer) and the CRD informer itself. +// - grpc is the Control Plane client. +// +// # When it runs +// +// cmd/api-server starts it only when config.ShouldPushClusterInventory +// agrees, which requires all of: +// +// - test triggers enabled. The picker is the only consumer, and the CRD +// informer needs cluster-scoped CRD read (the crd-reader ClusterRole in +// the testkube-api chart). Without that permission the informer logs a +// failed watch on every retry, so DISABLE_TEST_TRIGGERS turns it off +// rather than leave that noise on an install with nothing to show it to. +// - a Control Plane connection. A standalone agent serves discovery from +// its own API and has nowhere to push. +// - the listener capability. The Control Plane rejects pushes from any +// other agent, and a runner-only deployment has no CRD RBAC to watch with. +package inventory diff --git a/internal/sync/controller/doc.go b/internal/sync/controller/doc.go new file mode 100644 index 00000000000..1d3a7f13423 --- /dev/null +++ b/internal/sync/controller/doc.go @@ -0,0 +1,11 @@ +// Package controller contains the reconcilers that sync Kubernetes resources to the Control Plane store. +// +// Some store rejections do not clear on retry: an ownership conflict or an invalid resource (see +// syncagent.IsRejection). terminalOnRejection wraps such an error in reconcile.TerminalError, so +// controller-runtime stops the requeue of the resource. The reconcilers return every other error +// unchanged, and controller-runtime retries it with backoff. The reconcilers do not log these errors, +// because controller-runtime logs the returned error with the kind and the name of the resource. +// +// skipRejectedResource in cmd/api-server/superagentmigration.go applies the same rule during the +// SuperAgent migration. Without it, the migration retries a rejected resource forever. +package controller diff --git a/internal/sync/controller/errors.go b/internal/sync/controller/errors.go index e201b2d5e70..c74bb8a289c 100644 --- a/internal/sync/controller/errors.go +++ b/internal/sync/controller/errors.go @@ -1,24 +1,15 @@ package controller import ( - "errors" - "sigs.k8s.io/controller-runtime/pkg/reconcile" syncagent "github.com/kubeshop/testkube/internal/sync" ) -// terminalOnOwnershipConflict decides whether a failed store operation is worth retrying. -// -// An ownership conflict stands until somebody reassigns the owner or removes the resource from this -// agent's scope, so retrying only produces noise. Those are marked terminal to stop the requeue -// loop. Every other failure is returned unchanged so that it is retried with backoff. -// -// Nothing is logged here: controller-runtime logs whatever a reconciler returns, terminal or not, -// along with the kind and name of the resource, and the Control Plane's message naming the current -// owner travels along inside err. -func terminalOnOwnershipConflict(err error) error { - if errors.Is(err, syncagent.ErrOwnershipConflict) { +// terminalOnRejection marks a store rejection that no retry clears as terminal, and returns every +// other error unchanged. The package documentation gives the reasons. +func terminalOnRejection(err error) error { + if syncagent.IsRejection(err) { return reconcile.TerminalError(err) } return err diff --git a/internal/sync/controller/errors_test.go b/internal/sync/controller/errors_test.go index 40bcc7496af..56ca95b166e 100644 --- a/internal/sync/controller/errors_test.go +++ b/internal/sync/controller/errors_test.go @@ -10,33 +10,37 @@ import ( syncagent "github.com/kubeshop/testkube/internal/sync" ) -// An ownership conflict cannot be resolved by retrying, so it has to come back as a terminal error -// to stop controller-runtime requeueing it forever. -func TestOwnershipConflictIsTerminal(t *testing.T) { - err := fmt.Errorf("update TestWorkflow %q in store: %w", "smoke", syncagent.ErrOwnershipConflict) - - got := terminalOnOwnershipConflict(err) - - if !errors.Is(got, reconcile.TerminalError(nil)) { - t.Errorf("expected a terminal error so that controller-runtime stops requeueing, got %v", got) - } - // The Control Plane names the current owner in its message, and that message is the only clue - // the user gets from the agent log. - if !errors.Is(got, syncagent.ErrOwnershipConflict) { - t.Errorf("expected the ownership conflict to stay in the error chain, got %v", got) - } -} - -// Every other failure is transient as far as the agent can tell, so it has to stay retryable. -func TestOtherErrorsStayRetryable(t *testing.T) { - err := errors.New("connection refused") - - got := terminalOnOwnershipConflict(err) - - if errors.Is(got, reconcile.TerminalError(nil)) { - t.Errorf("expected a retryable error, got terminal error %v", got) +func TestTerminalOnRejection(t *testing.T) { + tests := []struct { + name string + err error + wantTerminal bool + }{ + { + name: "an ownership conflict stands until the owner changes", + err: fmt.Errorf("update TestWorkflow %q in store: %w", "smoke", syncagent.ErrOwnershipConflict), + wantTerminal: true, + }, + { + name: "an invalid resource stands until somebody changes it", + err: fmt.Errorf("update Webhook %q in store: %w", "hook", syncagent.ErrInvalidResource), + wantTerminal: true, + }, + { + name: "any other failure is transient as far as the agent can tell", + err: errors.New("connection refused"), + }, } - if !errors.Is(got, err) { - t.Errorf("expected the original error to be returned unchanged, got %v", got) + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + got := terminalOnRejection(tt.err) + + if errors.Is(got, reconcile.TerminalError(nil)) != tt.wantTerminal { + t.Errorf("terminal = %v, want %v, for %v", !tt.wantTerminal, tt.wantTerminal, got) + } + if !errors.Is(got, tt.err) { + t.Errorf("expected the original error to stay in the chain, because the agent log shows only its message, got %v", got) + } + }) } } diff --git a/internal/sync/controller/testtrigger.go b/internal/sync/controller/testtrigger.go index 85f4778e093..67276fb717f 100644 --- a/internal/sync/controller/testtrigger.go +++ b/internal/sync/controller/testtrigger.go @@ -36,7 +36,7 @@ func testTriggerSyncReconciler(client client.Reader, store TestTriggerStore) rec // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteTestTrigger(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete TestTrigger %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete TestTrigger %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil case err != nil: @@ -54,14 +54,14 @@ func testTriggerSyncReconciler(client client.Reader, store TestTriggerStore) rec // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteTestTrigger(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete TestTrigger %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete TestTrigger %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil } // Regular update so send the new object into the store. if err := store.UpdateOrCreateTestTrigger(ctx, trigger); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("update TestTrigger %q in store: %w", trigger.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("update TestTrigger %q in store: %w", trigger.Name, err)) } return ctrl.Result{}, nil diff --git a/internal/sync/controller/testworkflow.go b/internal/sync/controller/testworkflow.go index aa3acb1f5c9..47cae9a3823 100644 --- a/internal/sync/controller/testworkflow.go +++ b/internal/sync/controller/testworkflow.go @@ -36,7 +36,7 @@ func testWorkflowSyncReconciler(client client.Reader, store TestWorkflowStore) r // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteTestWorkflow(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete TestWorkflow %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete TestWorkflow %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil case err != nil: @@ -54,14 +54,14 @@ func testWorkflowSyncReconciler(client client.Reader, store TestWorkflowStore) r // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteTestWorkflow(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete TestWorkflow %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete TestWorkflow %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil } // Regular update so send the new object into the store. if err := store.UpdateOrCreateTestWorkflow(ctx, workflow); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("update TestWorkflow %q in store: %w", workflow.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("update TestWorkflow %q in store: %w", workflow.Name, err)) } return ctrl.Result{}, nil diff --git a/internal/sync/controller/testworkflowtemplate.go b/internal/sync/controller/testworkflowtemplate.go index 7e1aa8aafab..8b73d6b7d61 100644 --- a/internal/sync/controller/testworkflowtemplate.go +++ b/internal/sync/controller/testworkflowtemplate.go @@ -36,7 +36,7 @@ func testWorkflowTemplateSyncReconciler(client client.Reader, store TestWorkflow // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteTestWorkflowTemplate(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete TestWorkflowTemplate %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete TestWorkflowTemplate %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil case err != nil: @@ -54,14 +54,14 @@ func testWorkflowTemplateSyncReconciler(client client.Reader, store TestWorkflow // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteTestWorkflowTemplate(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete TestWorkflowTemplate %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete TestWorkflowTemplate %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil } // Regular update so send the new object into the store. if err := store.UpdateOrCreateTestWorkflowTemplate(ctx, workflowTemplate); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("update TestWorkflowTemplate %q in store: %w", workflowTemplate.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("update TestWorkflowTemplate %q in store: %w", workflowTemplate.Name, err)) } return ctrl.Result{}, nil diff --git a/internal/sync/controller/webhook.go b/internal/sync/controller/webhook.go index 3a53f1db49c..29f3054e0fb 100644 --- a/internal/sync/controller/webhook.go +++ b/internal/sync/controller/webhook.go @@ -36,7 +36,7 @@ func webhookSyncReconciler(client client.Reader, store WebhookStore) reconcile.R // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteWebhook(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete Webhook %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete Webhook %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil case err != nil: @@ -54,14 +54,14 @@ func webhookSyncReconciler(client client.Reader, store WebhookStore) reconcile.R // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteWebhook(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete Webhook %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete Webhook %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil } // Regular update so send the new object into the store. if err := store.UpdateOrCreateWebhook(ctx, webhook); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("update Webhook %q in store: %w", webhook.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("update Webhook %q in store: %w", webhook.Name, err)) } return ctrl.Result{}, nil diff --git a/internal/sync/controller/webhooktemplate.go b/internal/sync/controller/webhooktemplate.go index fcc103bcd63..2d2dacbe8bd 100644 --- a/internal/sync/controller/webhooktemplate.go +++ b/internal/sync/controller/webhooktemplate.go @@ -36,7 +36,7 @@ func webhookTemplateSyncReconciler(client client.Reader, store WebhookTemplateSt // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteWebhookTemplate(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete WebhookTemplate %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete WebhookTemplate %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil case err != nil: @@ -54,14 +54,14 @@ func webhookTemplateSyncReconciler(client client.Reader, store WebhookTemplateSt // Passing the name here rather than the namespaced name as generally we refer to objects // purely by their name. if err := store.DeleteWebhookTemplate(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete WebhookTemplate %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete WebhookTemplate %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil } // Regular update so send the new object into the store. if err := store.UpdateOrCreateWebhookTemplate(ctx, template); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("update WebhookTemplate %q in store: %w", template.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("update WebhookTemplate %q in store: %w", template.Name, err)) } return ctrl.Result{}, nil diff --git a/internal/sync/controller/workflow_trigger.go b/internal/sync/controller/workflow_trigger.go index 14de06b105a..aa3bced2e96 100644 --- a/internal/sync/controller/workflow_trigger.go +++ b/internal/sync/controller/workflow_trigger.go @@ -33,7 +33,7 @@ func workflowTriggerSyncReconciler(client client.Reader, store WorkflowTriggerSt switch { case errors.IsNotFound(err): if err := store.DeleteWorkflowTrigger(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete WorkflowTrigger %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete WorkflowTrigger %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil case err != nil: @@ -42,13 +42,13 @@ func workflowTriggerSyncReconciler(client client.Reader, store WorkflowTriggerSt if !trigger.DeletionTimestamp.IsZero() { if err := store.DeleteWorkflowTrigger(ctx, req.Name); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("delete WorkflowTrigger %q from store: %w", req.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("delete WorkflowTrigger %q from store: %w", req.Name, err)) } return ctrl.Result{}, nil } if err := store.UpdateOrCreateWorkflowTrigger(ctx, trigger); err != nil { - return ctrl.Result{}, terminalOnOwnershipConflict(fmt.Errorf("update WorkflowTrigger %q in store: %w", trigger.Name, err)) + return ctrl.Result{}, terminalOnRejection(fmt.Errorf("update WorkflowTrigger %q in store: %w", trigger.Name, err)) } return ctrl.Result{}, nil diff --git a/internal/sync/errors.go b/internal/sync/errors.go index f78dedbbfd7..cb15f9f43c1 100644 --- a/internal/sync/errors.go +++ b/internal/sync/errors.go @@ -9,3 +9,14 @@ import "errors" // Ownership is reassigned by setting the testkube.io/gitops-owner annotation on the Kubernetes // resource, or by removing the resource from the owning agent's scope. var ErrOwnershipConflict = errors.New("resource is owned by another GitOps agent") + +// ErrInvalidResource is returned by a store when the Control Plane rejects a resource as invalid, +// for example a webhook condition with an unknown type. The Control Plane rejects the same resource +// on every retry, so callers treat it as terminal until somebody changes the resource. +var ErrInvalidResource = errors.New("the Control Plane rejected the resource as invalid") + +// IsRejection reports whether the Control Plane rejected a resource in a way that no retry clears: +// an ownership conflict or an invalid resource. +func IsRejection(err error) bool { + return errors.Is(err, ErrOwnershipConflict) || errors.Is(err, ErrInvalidResource) +} diff --git a/internal/sync/errors_test.go b/internal/sync/errors_test.go new file mode 100644 index 00000000000..29cd5c04bc9 --- /dev/null +++ b/internal/sync/errors_test.go @@ -0,0 +1,26 @@ +package sync + +import ( + "errors" + "fmt" + "testing" +) + +func TestIsRejection(t *testing.T) { + tests := []struct { + name string + err error + want bool + }{ + {name: "an ownership conflict never clears by retry", err: fmt.Errorf("sync: %w", ErrOwnershipConflict), want: true}, + {name: "an invalid resource never clears by retry", err: fmt.Errorf("sync: %w", ErrInvalidResource), want: true}, + {name: "any other failure is retried", err: errors.New("connection refused")}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + if got := IsRejection(tt.err); got != tt.want { + t.Errorf("IsRejection() = %v, want %v", got, tt.want) + } + }) + } +} diff --git a/internal/sync/grpc/errors.go b/internal/sync/grpc/errors.go index 85e35e9a81f..da501d69fc7 100644 --- a/internal/sync/grpc/errors.go +++ b/internal/sync/grpc/errors.go @@ -13,15 +13,19 @@ import ( // sync package, so that callers can react to the outcome without depending on gRPC. // // The Control Plane only uses FailedPrecondition on the sync API to reject a change that conflicts -// with another GitOps agent's ownership of the resource. The original error is kept in the chain -// because its message names the current owner. +// with another GitOps agent's ownership of the resource, and InvalidArgument to reject a resource +// that it cannot store. The result keeps the original error in the chain, because its message names +// the current owner or the invalid value. func translateError(err error) error { if err == nil { return nil } - if status.Code(err) == codes.FailedPrecondition { + switch status.Code(err) { + case codes.FailedPrecondition: return fmt.Errorf("%w: %w", syncagent.ErrOwnershipConflict, err) + case codes.InvalidArgument: + return fmt.Errorf("%w: %w", syncagent.ErrInvalidResource, err) } return err diff --git a/internal/sync/grpc/errors_test.go b/internal/sync/grpc/errors_test.go index 4a12d619bc4..da65ed07dd6 100644 --- a/internal/sync/grpc/errors_test.go +++ b/internal/sync/grpc/errors_test.go @@ -8,56 +8,58 @@ import ( "google.golang.org/grpc/codes" "google.golang.org/grpc/status" + executorv1 "github.com/kubeshop/testkube/api/executor/v1" testworkflowsv1 "github.com/kubeshop/testkube/api/testworkflows/v1" syncagent "github.com/kubeshop/testkube/internal/sync" + syncgrpc "github.com/kubeshop/testkube/internal/sync/grpc" ) -// The Control Plane signals an ownership conflict with FailedPrecondition. Callers act on the -// sentinel rather than on gRPC codes, so the client has to translate it. -func TestOwnershipConflictIsRecognised(t *testing.T) { - const rejection = `TestWorkflow "smoke" is owned by GitOps agent "team-a-gitops"` - srv := testSrv{Err: status.Error(codes.FailedPrecondition, rejection)} - client := startGRPCTestConnection(t, &srv) - - tests := map[string]func() error{ - "update": func() error { - return client.UpdateOrCreateTestWorkflow(t.Context(), testworkflowsv1.TestWorkflow{}) - }, - "delete": func() error { - return client.DeleteTestWorkflow(t.Context(), "smoke") - }, +func TestTranslateError(t *testing.T) { + updateWorkflow := func(t *testing.T, c syncgrpc.Client) error { + return c.UpdateOrCreateTestWorkflow(t.Context(), testworkflowsv1.TestWorkflow{}) } - - for name, call := range tests { - t.Run(name, func(t *testing.T) { - err := call() - - if !errors.Is(err, syncagent.ErrOwnershipConflict) { - t.Errorf("expected an ownership conflict, got %v", err) - } - // The Control Plane's message names the owner, so it has to survive translation. - if !strings.Contains(err.Error(), rejection) { - t.Errorf("expected the Control Plane's message to be preserved, got %v", err) - } - }) + deleteWorkflow := func(t *testing.T, c syncgrpc.Client) error { + return c.DeleteTestWorkflow(t.Context(), "smoke") + } + updateWebhook := func(t *testing.T, c syncgrpc.Client) error { + return c.UpdateOrCreateWebhook(t.Context(), executorv1.Webhook{}) } -} -// Any other code is a transient or unrelated failure and must not be mistaken for a conflict, so -// that the reconciler keeps retrying it. -func TestOtherStatusCodesAreNotConflicts(t *testing.T) { - for _, code := range []codes.Code{codes.Internal, codes.Unavailable, codes.InvalidArgument, codes.NotFound} { - t.Run(code.String(), func(t *testing.T) { - srv := testSrv{Err: status.Error(code, "boom")} - client := startGRPCTestConnection(t, &srv) + tests := []struct { + name string + code codes.Code + call func(*testing.T, syncgrpc.Client) error + // want is the sentinel that the error wraps, or nil for an error that stays retryable. + want error + }{ + {name: "FailedPrecondition on update is an ownership conflict", code: codes.FailedPrecondition, call: updateWorkflow, want: syncagent.ErrOwnershipConflict}, + {name: "FailedPrecondition on delete is an ownership conflict", code: codes.FailedPrecondition, call: deleteWorkflow, want: syncagent.ErrOwnershipConflict}, + {name: "InvalidArgument is an invalid resource", code: codes.InvalidArgument, call: updateWebhook, want: syncagent.ErrInvalidResource}, + {name: "Internal stays retryable", code: codes.Internal, call: updateWorkflow}, + {name: "Unavailable stays retryable", code: codes.Unavailable, call: updateWorkflow}, + {name: "NotFound stays retryable", code: codes.NotFound, call: updateWorkflow}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + const message = "message of the Control Plane" + client := startGRPCTestConnection(t, &testSrv{Err: status.Error(tt.code, message)}) - err := client.UpdateOrCreateTestWorkflow(t.Context(), testworkflowsv1.TestWorkflow{}) + err := tt.call(t, client) if err == nil { t.Fatal("expected an error") } - if errors.Is(err, syncagent.ErrOwnershipConflict) { - t.Errorf("%s must not be treated as an ownership conflict, got %v", code, err) + if tt.want == nil { + if syncagent.IsRejection(err) { + t.Errorf("expected a retryable error, got rejection %v", err) + } + return + } + if !errors.Is(err, tt.want) { + t.Errorf("expected %v in the chain, got %v", tt.want, err) + } + if !strings.Contains(err.Error(), message) { + t.Errorf("expected the message of the Control Plane to stay in the error, because it names the owner or the invalid value, got %v", err) } }) } diff --git a/k8s/crd/executor.testkube.io_webhooks.yaml b/k8s/crd/executor.testkube.io_webhooks.yaml index b524fe4b504..b66c6ab453e 100644 --- a/k8s/crd/executor.testkube.io_webhooks.yaml +++ b/k8s/crd/executor.testkube.io_webhooks.yaml @@ -101,6 +101,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed diff --git a/k8s/crd/executor.testkube.io_webhooktemplates.yaml b/k8s/crd/executor.testkube.io_webhooktemplates.yaml index 50a1a05e3d3..d73c1870b5a 100644 --- a/k8s/crd/executor.testkube.io_webhooktemplates.yaml +++ b/k8s/crd/executor.testkube.io_webhooktemplates.yaml @@ -102,6 +102,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed diff --git a/k8s/helm/testkube-crds/templates/_generated_crds.tpl b/k8s/helm/testkube-crds/templates/_generated_crds.tpl index 974e5d50e6c..723b08b0ae8 100644 --- a/k8s/helm/testkube-crds/templates/_generated_crds.tpl +++ b/k8s/helm/testkube-crds/templates/_generated_crds.tpl @@ -291,6 +291,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed @@ -504,6 +507,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed diff --git a/k8s/helm/testkube-operator/charts/testkube-crds/templates/_generated_crds.tpl b/k8s/helm/testkube-operator/charts/testkube-crds/templates/_generated_crds.tpl index 974e5d50e6c..723b08b0ae8 100644 --- a/k8s/helm/testkube-operator/charts/testkube-crds/templates/_generated_crds.tpl +++ b/k8s/helm/testkube-operator/charts/testkube-crds/templates/_generated_crds.tpl @@ -291,6 +291,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed @@ -504,6 +507,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed diff --git a/k8s/helm/testkube-operator/values.yaml b/k8s/helm/testkube-operator/values.yaml index b042ba85754..91218bf5779 100644 --- a/k8s/helm/testkube-operator/values.yaml +++ b/k8s/helm/testkube-operator/values.yaml @@ -82,7 +82,7 @@ proxy: # -- Testkube Operator rbac-proxy image repository repository: brancz/kube-rbac-proxy # -- Testkube Operator rbac-proxy image tag - tag: "v0.22.1" + tag: "v0.23.0" # -- Testkube Operator rbac-proxy k8s secret for private registries pullSecrets: [] # -- Testkube Operator rbac-proxy resource settings diff --git a/k8s/helm/testkube-runner/charts/testkube-crds/templates/_generated_crds.tpl b/k8s/helm/testkube-runner/charts/testkube-crds/templates/_generated_crds.tpl index 974e5d50e6c..723b08b0ae8 100644 --- a/k8s/helm/testkube-runner/charts/testkube-crds/templates/_generated_crds.tpl +++ b/k8s/helm/testkube-runner/charts/testkube-crds/templates/_generated_crds.tpl @@ -291,6 +291,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed @@ -504,6 +507,9 @@ spec: - end-testworkflow-aborted - end-testworkflow-canceled - end-testworkflow-not-passed + - end-testworkflow-test-failure + - end-testworkflow-infrastructure-failure + - end-testworkflow-configuration-error - become-testworkflow-up - become-testworkflow-down - become-testworkflow-failed diff --git a/k8s/helm/testkube/Chart.lock b/k8s/helm/testkube/Chart.lock index 1d79848a6a3..79da1ef04fd 100644 --- a/k8s/helm/testkube/Chart.lock +++ b/k8s/helm/testkube/Chart.lock @@ -10,7 +10,7 @@ dependencies: version: 18.6-0 - name: cloudnative-pg repository: https://cloudnative-pg.github.io/charts - version: 0.29.0 + version: 0.29.1 - name: cluster repository: https://cloudnative-pg.github.io/charts version: 0.8.1 @@ -23,5 +23,5 @@ dependencies: - name: global repository: file://../global version: 0.1.3 -digest: sha256:c4e2cc0eb7c2dff4841282bde24f65530e94971187636e39d831216e54203c1b -generated: "2026-09-01T14:33:38.572893852Z" +digest: sha256:81b85d4bd6dde573d94050376d6c445baf6ec40e2d7ab2b5172a7d025980280e +generated: "2026-09-24T04:12:57.849380098Z" diff --git a/k8s/helm/testkube/Chart.yaml b/k8s/helm/testkube/Chart.yaml index f14b916dbbd..8262957c58c 100644 --- a/k8s/helm/testkube/Chart.yaml +++ b/k8s/helm/testkube/Chart.yaml @@ -17,7 +17,7 @@ dependencies: version: 18.6-0 repository: oci://us-east1-docker.pkg.dev/testkube-cloud-372110/testkube - name: cloudnative-pg - version: 0.29.0 + version: 0.29.1 repository: https://cloudnative-pg.github.io/charts condition: cloudnative-pg.enabled - name: cluster diff --git a/k8s/helm/testkube/values.yaml b/k8s/helm/testkube/values.yaml index 83d18b7234b..af54bfa2b5c 100644 --- a/k8s/helm/testkube/values.yaml +++ b/k8s/helm/testkube/values.yaml @@ -1315,7 +1315,7 @@ testkube-operator: # -- Testkube Operator rbac-proxy image repository repository: brancz/kube-rbac-proxy # -- Testkube Operator rbac-proxy image tag - tag: "v0.22.1" + tag: "v0.23.0" # -- Testkube Operator rbac-proxy k8s secret for private registries pullSecrets: [] # -- Testkube Operator rbac-proxy resource settings diff --git a/pkg/api/v1/testkube/model_event_extended.go b/pkg/api/v1/testkube/model_event_extended.go index 2853b0984ad..c31e8375371 100644 --- a/pkg/api/v1/testkube/model_event_extended.go +++ b/pkg/api/v1/testkube/model_event_extended.go @@ -179,6 +179,14 @@ func (e Event) Valid(groupId, selector string, types []EventType) (matchedTypes typesMatch := false for _, t := range types { + if detailsType, ok := t.causeStatusDetailsType(); ok { + if e.hasFailureCause(detailsType) { + typesMatch = true + matchedTypes = append(matchedTypes, t) + } + continue + } + ts := []EventType{t} if t.IsBecome() { ts = t.MapBecomeToRegular() @@ -210,6 +218,16 @@ func (e Event) Valid(groupId, selector string, types []EventType) (matchedTypes return } +// hasFailureCause reports whether the event is the not-passed end event of an execution whose +// status details carry the type. +func (e Event) hasFailureCause(detailsType StatusDetailsType) bool { + if e.Type() != END_TESTWORKFLOW_NOT_PASSED_EventType || e.TestWorkflowExecution == nil { + return false + } + result := e.TestWorkflowExecution.Result + return result != nil && result.StatusDetails != nil && StatusDetailsType(result.StatusDetails.Type_) == detailsType +} + // GetResourceId implmenents generic event trigger func (e Event) GetResourceId() string { return e.ResourceId diff --git a/pkg/api/v1/testkube/model_event_extended_test.go b/pkg/api/v1/testkube/model_event_extended_test.go index 5ccf48d8f14..d49a56ec5ef 100644 --- a/pkg/api/v1/testkube/model_event_extended_test.go +++ b/pkg/api/v1/testkube/model_event_extended_test.go @@ -134,3 +134,104 @@ func TestNewWorkflowEvents_UseExplicitRoutingGroup(t *testing.T) { assert.Equal(t, "env-123", queueEvent.GroupId) assert.Equal(t, "env-123", startEvent.GroupId) } + +func TestEvent_Valid_CauseEvents(t *testing.T) { + withDetails := func(detailsType StatusDetailsType) *TestWorkflowExecution { + return &TestWorkflowExecution{ + Workflow: &TestWorkflow{Labels: map[string]string{"team": "platform"}}, + Result: &TestWorkflowResult{StatusDetails: &TestWorkflowStatusDetails{Type_: string(detailsType)}}, + } + } + + tests := []struct { + name string + eventType *EventType + execution *TestWorkflowExecution + selector string + types []EventType + wantTypes []EventType + wantValid bool + }{ + { + name: "the infrastructure event matches the not-passed event of an execution failure", + eventType: EventEndTestWorkflowNotPassed, + execution: withDetails(StatusDetailsTypeExecutionFailure), + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + wantTypes: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + wantValid: true, + }, + { + name: "the test failure event matches the not-passed event of a step failure", + eventType: EventEndTestWorkflowNotPassed, + execution: withDetails(StatusDetailsTypeStepFailure), + types: []EventType{END_TESTWORKFLOW_TEST_FAILURE_EventType}, + wantTypes: []EventType{END_TESTWORKFLOW_TEST_FAILURE_EventType}, + wantValid: true, + }, + { + name: "the configuration error event matches the not-passed event of an init failure", + eventType: EventEndTestWorkflowNotPassed, + execution: withDetails(StatusDetailsTypeInitFailure), + types: []EventType{END_TESTWORKFLOW_CONFIGURATION_ERROR_EventType}, + wantTypes: []EventType{END_TESTWORKFLOW_CONFIGURATION_ERROR_EventType}, + wantValid: true, + }, + { + name: "the failed event of the same execution does not match, so a listener gets one call", + eventType: EventEndTestWorkflowFailed, + execution: withDetails(StatusDetailsTypeExecutionFailure), + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + { + name: "the aborted event of the same execution does not match", + eventType: EventEndTestWorkflowAborted, + execution: withDetails(StatusDetailsTypeExecutionFailure), + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + { + name: "another type does not match", + eventType: EventEndTestWorkflowNotPassed, + execution: withDetails(StatusDetailsTypeStepFailure), + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + { + name: "an execution without status details does not match", + eventType: EventEndTestWorkflowNotPassed, + execution: &TestWorkflowExecution{Result: &TestWorkflowResult{}}, + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + { + name: "an execution without a result does not match", + eventType: EventEndTestWorkflowNotPassed, + execution: &TestWorkflowExecution{}, + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + { + name: "a cause event next to the not-passed event matches both on one event", + eventType: EventEndTestWorkflowNotPassed, + execution: withDetails(StatusDetailsTypeExecutionFailure), + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType, END_TESTWORKFLOW_NOT_PASSED_EventType}, + wantTypes: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType, END_TESTWORKFLOW_NOT_PASSED_EventType}, + wantValid: true, + }, + { + name: "the selector still applies to a cause event", + eventType: EventEndTestWorkflowNotPassed, + execution: withDetails(StatusDetailsTypeExecutionFailure), + selector: "team=qa", + types: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + wantTypes: []EventType{END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + } + + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + e := Event{Type_: tt.eventType, TestWorkflowExecution: tt.execution} + + types, valid := e.Valid("", tt.selector, tt.types) + + assert.Equal(t, tt.wantTypes, types) + assert.Equal(t, tt.wantValid, valid) + }) + } +} diff --git a/pkg/api/v1/testkube/model_event_type.go b/pkg/api/v1/testkube/model_event_type.go index ff90459f07a..3f2abfe6c1c 100644 --- a/pkg/api/v1/testkube/model_event_type.go +++ b/pkg/api/v1/testkube/model_event_type.go @@ -13,20 +13,23 @@ type EventType string // List of EventType const ( - QUEUE_TESTWORKFLOW_EventType EventType = "queue-testworkflow" - START_TESTWORKFLOW_EventType EventType = "start-testworkflow" - END_TESTWORKFLOW_SUCCESS_EventType EventType = "end-testworkflow-success" - END_TESTWORKFLOW_FAILED_EventType EventType = "end-testworkflow-failed" - END_TESTWORKFLOW_ABORTED_EventType EventType = "end-testworkflow-aborted" - END_TESTWORKFLOW_CANCELED_EventType EventType = "end-testworkflow-canceled" - END_TESTWORKFLOW_NOT_PASSED_EventType EventType = "end-testworkflow-not-passed" - BECOME_TESTWORKFLOW_UP_EventType EventType = "become-testworkflow-up" - BECOME_TESTWORKFLOW_DOWN_EventType EventType = "become-testworkflow-down" - BECOME_TESTWORKFLOW_FAILED_EventType EventType = "become-testworkflow-failed" - BECOME_TESTWORKFLOW_ABORTED_EventType EventType = "become-testworkflow-aborted" - BECOME_TESTWORKFLOW_CANCELED_EventType EventType = "become-testworkflow-canceled" - BECOME_TESTWORKFLOW_NOT_PASSED_EventType EventType = "become-testworkflow-not-passed" - CREATED_EventType EventType = "created" - UPDATED_EventType EventType = "updated" - DELETED_EventType EventType = "deleted" + QUEUE_TESTWORKFLOW_EventType EventType = "queue-testworkflow" + START_TESTWORKFLOW_EventType EventType = "start-testworkflow" + END_TESTWORKFLOW_SUCCESS_EventType EventType = "end-testworkflow-success" + END_TESTWORKFLOW_FAILED_EventType EventType = "end-testworkflow-failed" + END_TESTWORKFLOW_ABORTED_EventType EventType = "end-testworkflow-aborted" + END_TESTWORKFLOW_CANCELED_EventType EventType = "end-testworkflow-canceled" + END_TESTWORKFLOW_NOT_PASSED_EventType EventType = "end-testworkflow-not-passed" + END_TESTWORKFLOW_TEST_FAILURE_EventType EventType = "end-testworkflow-test-failure" + END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType EventType = "end-testworkflow-infrastructure-failure" + END_TESTWORKFLOW_CONFIGURATION_ERROR_EventType EventType = "end-testworkflow-configuration-error" + BECOME_TESTWORKFLOW_UP_EventType EventType = "become-testworkflow-up" + BECOME_TESTWORKFLOW_DOWN_EventType EventType = "become-testworkflow-down" + BECOME_TESTWORKFLOW_FAILED_EventType EventType = "become-testworkflow-failed" + BECOME_TESTWORKFLOW_ABORTED_EventType EventType = "become-testworkflow-aborted" + BECOME_TESTWORKFLOW_CANCELED_EventType EventType = "become-testworkflow-canceled" + BECOME_TESTWORKFLOW_NOT_PASSED_EventType EventType = "become-testworkflow-not-passed" + CREATED_EventType EventType = "created" + UPDATED_EventType EventType = "updated" + DELETED_EventType EventType = "deleted" ) diff --git a/pkg/api/v1/testkube/model_event_type_extended.go b/pkg/api/v1/testkube/model_event_type_extended.go index 981d2389fc0..3d2f772c8d2 100644 --- a/pkg/api/v1/testkube/model_event_type_extended.go +++ b/pkg/api/v1/testkube/model_event_type_extended.go @@ -33,6 +33,23 @@ var ( EventUpdated = EventTypePtr(UPDATED_EventType) ) +// causeEventStatusDetailsTypes maps each event for the cause of a failure to its status +// details type. Nothing emits these events. Each one matches the not-passed end event of an +// execution with that type. An execution that does not pass emits exactly one not-passed +// event, so each cause event gives at most one call per execution. +var causeEventStatusDetailsTypes = map[EventType]StatusDetailsType{ + END_TESTWORKFLOW_TEST_FAILURE_EventType: StatusDetailsTypeStepFailure, + END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType: StatusDetailsTypeExecutionFailure, + END_TESTWORKFLOW_CONFIGURATION_ERROR_EventType: StatusDetailsTypeInitFailure, +} + +// causeStatusDetailsType returns the status details type of an event for the cause of a +// failure, and false for every other event type. +func (t EventType) causeStatusDetailsType() (StatusDetailsType, bool) { + detailsType, ok := causeEventStatusDetailsTypes[t] + return detailsType, ok +} + func (t EventType) IsBecome() bool { types := []EventType{ BECOME_TESTWORKFLOW_UP_EventType, diff --git a/pkg/api/v1/testkube/status_classify.go b/pkg/api/v1/testkube/status_classify.go index 4a577ecfb08..660c1081d53 100644 --- a/pkg/api/v1/testkube/status_classify.go +++ b/pkg/api/v1/testkube/status_classify.go @@ -22,12 +22,8 @@ func (r *TestWorkflowResult) ClassifyStatus(sigSequence []TestWorkflowSignature, return nil } - if stop.Actor == StopActorUser || stop.Actor == StopActorAPI { - reason := stop.Reason - if reason == "" { - reason = StopReasonUserCancel - } - return r.stopDetails(sigSequence, stop.Actor, reason) + if stop.Actor.IsPerson() { + return r.stopDetails(sigSequence, stop.Actor, stop.Reason) } reason := stop.Reason diff --git a/pkg/api/v1/testkube/status_details.go b/pkg/api/v1/testkube/status_details.go index e295d84fe10..53ec5245370 100644 --- a/pkg/api/v1/testkube/status_details.go +++ b/pkg/api/v1/testkube/status_details.go @@ -73,10 +73,9 @@ var statusDetailsTypes = map[string]StatusDetailsType{ // StatusDetailsTypeOf returns the layer for a stop. A person as the actor gives // user-cancel for every reason, because a cancel is never a failure. A person who stops -// all executions of a workflow makes the same decision as a person who cancels one. The -// abort endpoint of the standalone agent acts for the person who calls it. +// all executions of a workflow makes the same decision as a person who cancels one. func StatusDetailsTypeOf(actor StopActor, reason string) StatusDetailsType { - if actor == StopActorUser || actor == StopActorAPI { + if actor.IsPerson() { return StatusDetailsTypeUserCancel } if detailsType, ok := statusDetailsTypes[reason]; ok { @@ -86,11 +85,14 @@ func StatusDetailsTypeOf(actor StopActor, reason string) StatusDetailsType { } // NewStatusDetails builds the object for a stop. The step and the message identify the -// step that holds the cause, and both stay empty when no step holds one. An empty reason -// becomes unknown, because the field is required and a reader must always get a code. +// step that holds the cause, and both stay empty when no step holds one. The reason is +// required, so an empty reason becomes user-cancel for a person and unknown for other actors. func NewStatusDetails(actor StopActor, reason, step, message string) *TestWorkflowStatusDetails { if reason == "" { reason = string(StopReasonUnknown) + if actor.IsPerson() { + reason = string(StopReasonUserCancel) + } } return &TestWorkflowStatusDetails{ Type_: string(StatusDetailsTypeOf(actor, reason)), diff --git a/pkg/api/v1/testkube/status_details_test.go b/pkg/api/v1/testkube/status_details_test.go index faf9ed710a1..00ef3504d72 100644 --- a/pkg/api/v1/testkube/status_details_test.go +++ b/pkg/api/v1/testkube/status_details_test.go @@ -120,12 +120,46 @@ func TestNewStatusDetails(t *testing.T) { }, }, { - name: "a cancel by a person carries no reason", - actor: StopActorUser, - reason: "", + name: "a cancel by a person without a reason", + actor: StopActorUser, + want: TestWorkflowStatusDetails{ + Type_: string(StatusDetailsTypeUserCancel), + Reason: string(StopReasonUserCancel), + Actor: string(StopActorUser), + }, + }, + { + name: "a cancel through the API without a reason", + actor: StopActorAPI, want: TestWorkflowStatusDetails{ Type_: string(StatusDetailsTypeUserCancel), + Reason: string(StopReasonUserCancel), + Actor: string(StopActorAPI), + }, + }, + { + name: "a control plane stop without a reason", + actor: StopActorControlPlane, + want: TestWorkflowStatusDetails{ + Type_: string(StatusDetailsTypeUnknown), Reason: string(StopReasonUnknown), + Actor: string(StopActorControlPlane), + }, + }, + { + name: "a stop without an actor and without a reason", + want: TestWorkflowStatusDetails{ + Type_: string(StatusDetailsTypeUnknown), + Reason: string(StopReasonUnknown), + }, + }, + { + name: "a cancel by a person keeps the reason it carries", + actor: StopActorUser, + reason: string(StopReasonAbortAll), + want: TestWorkflowStatusDetails{ + Type_: string(StatusDetailsTypeUserCancel), + Reason: string(StopReasonAbortAll), Actor: string(StopActorUser), }, }, diff --git a/pkg/api/v1/testkube/stop_actor.go b/pkg/api/v1/testkube/stop_actor.go index eadd9e8ba62..8b0e5f7ba1f 100644 --- a/pkg/api/v1/testkube/stop_actor.go +++ b/pkg/api/v1/testkube/stop_actor.go @@ -24,6 +24,12 @@ const ( StopActorSystem StopActor = "system" ) +// IsPerson reports whether a person decided the stop. The abort endpoint of the +// standalone agent acts for the person who calls it, so it counts as a person too. +func (a StopActor) IsPerson() bool { + return a == StopActorUser || a == StopActorAPI +} + // Sentence returns the words the result reader uses for the actor after // "The execution has been aborted". The reason, when set, follows them. func (a StopActor) Sentence() string { diff --git a/pkg/api/v1/testkube/stop_actor_test.go b/pkg/api/v1/testkube/stop_actor_test.go index c007daa0f0a..3078b4d6bfc 100644 --- a/pkg/api/v1/testkube/stop_actor_test.go +++ b/pkg/api/v1/testkube/stop_actor_test.go @@ -12,13 +12,13 @@ func TestStopActor_Sentence(t *testing.T) { actor StopActor want string }{ - {name: "user", actor: StopActorUser, want: "by the user"}, - {name: "control plane", actor: StopActorControlPlane, want: "by the control plane"}, + {name: "a user is a person", actor: StopActorUser, want: "by the user"}, + {name: "the control plane is not a person", actor: StopActorControlPlane, want: "by the control plane"}, {name: "trigger", actor: StopActorTrigger, want: "because its trigger was deleted"}, {name: "fail-fast", actor: StopActorFailFast, want: "because another parallel worker failed"}, {name: "runner", actor: StopActorRunner, want: "by the runner"}, {name: "quality loop", actor: StopActorQualityLoop, want: "by the quality loop"}, - {name: "api", actor: StopActorAPI, want: "through the API"}, + {name: "the abort endpoint acts for a person", actor: StopActorAPI, want: "through the API"}, {name: "system", actor: StopActorSystem, want: "by the system"}, {name: "unknown value falls back to the system", actor: StopActor("later-added"), want: "by the system"}, } @@ -28,3 +28,21 @@ func TestStopActor_Sentence(t *testing.T) { }) } } + +func TestStopActor_IsPerson(t *testing.T) { + tests := []struct { + name string + actor StopActor + want bool + }{ + {name: "a user is a person", actor: StopActorUser, want: true}, + {name: "the abort endpoint acts for a person", actor: StopActorAPI, want: true}, + {name: "the control plane is not a person", actor: StopActorControlPlane, want: false}, + {name: "a stop without an actor has no person", actor: "", want: false}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + assert.Equal(t, tt.want, tt.actor.IsPerson()) + }) + } +} diff --git a/pkg/event/emitter_test.go b/pkg/event/emitter_test.go index 20bb6e2ecfb..b782d951ec2 100644 --- a/pkg/event/emitter_test.go +++ b/pkg/event/emitter_test.go @@ -921,3 +921,80 @@ func TestEmitterReconcileInterval(t *testing.T) { assert.Equal(t, DefaultReconcileInterval, emitter.reconcileInterval) }) } + +// TestEmitter_CauseEvents checks the delivery of the cause events: each one arrives with its own +// type, and every subscribed type that matches gets its own call, as for the become events. +func TestEmitter_CauseEvents(t *testing.T) { + eventBus := bus.NewEventBusMock() + mockCtrl := gomock.NewController(t) + t.Cleanup(mockCtrl.Finish) + mockLeaseRepository := leasebackend.NewMockRepository(mockCtrl) + mockLeaseRepository.EXPECT().TryAcquire(gomock.Any(), gomock.Any(), gomock.Any()).Return(true, nil).AnyTimes() + emitter := NewEmitter(eventBus, mockLeaseRepository, "agentevents", "", DefaultEventTTL, DefaultEventCacheCapacity) + t.Cleanup(emitter.eventCache.Stop) + + tests := []struct { + name string + types []testkube.EventType + wantTypes []testkube.EventType + }{ + { + name: "a cause event arrives with its own type", + types: []testkube.EventType{testkube.END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + wantTypes: []testkube.EventType{testkube.END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + { + name: "a cause event next to not-passed gives one call for each", + types: []testkube.EventType{testkube.END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType, testkube.END_TESTWORKFLOW_NOT_PASSED_EventType}, + wantTypes: []testkube.EventType{testkube.END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType, testkube.END_TESTWORKFLOW_NOT_PASSED_EventType}, + }, + { + name: "a status event next to a cause event gives one call for each", + types: []testkube.EventType{testkube.END_TESTWORKFLOW_FAILED_EventType, testkube.END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + wantTypes: []testkube.EventType{testkube.END_TESTWORKFLOW_FAILED_EventType, testkube.END_TESTWORKFLOW_INFRASTRUCTURE_FAILURE_EventType}, + }, + { + name: "another cause gets no call", + types: []testkube.EventType{testkube.END_TESTWORKFLOW_TEST_FAILURE_EventType}, + }, + } + + listeners := make([]*dummy.DummyListener, len(tests)) + for i, tt := range tests { + listeners[i] = &dummy.DummyListener{Id: tt.name, Types: tt.types} + emitter.Register(listeners[i]) + } + + ctx, cancel := context.WithCancel(t.Context()) + t.Cleanup(cancel) + go emitter.Listen(ctx) + time.Sleep(50 * time.Millisecond) + + // A failed execution emits its status event and one not-passed event. + execution := &testkube.TestWorkflowExecution{ + Id: "exec-oom", + Workflow: &testkube.TestWorkflow{Name: "tw1"}, + Result: &testkube.TestWorkflowResult{ + StatusDetails: &testkube.TestWorkflowStatusDetails{Type_: string(testkube.StatusDetailsTypeExecutionFailure)}, + }, + } + emitter.Notify(testkube.Event{Id: "event-failed", Type_: testkube.EventEndTestWorkflowFailed, TestWorkflowExecution: execution}) + emitter.Notify(testkube.Event{Id: "event-not-passed", Type_: testkube.EventEndTestWorkflowNotPassed, TestWorkflowExecution: execution}) + + assert.Eventually(t, func() bool { + for i, tt := range tests { + if listeners[i].GetNotificationCount() < len(tt.wantTypes) { + return false + } + } + return true + }, 5*time.Second, 20*time.Millisecond) + // Give an unexpected extra call the time to arrive. + time.Sleep(100 * time.Millisecond) + + for i, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + assert.ElementsMatch(t, tt.wantTypes, listeners[i].GetReceivedEventTypes()) + }) + } +} diff --git a/pkg/mcp/README.md b/pkg/mcp/README.md index cfa383d209e..2e43a94f898 100644 --- a/pkg/mcp/README.md +++ b/pkg/mcp/README.md @@ -47,7 +47,7 @@ This flexibility allows the same MCP tools to work in different deployment scena ### Available Tools -The MCP server exposes up to 34 tools organized into the categories below. The two +The MCP server exposes up to 43 tools organized into the categories below. The two Query tools register conditionally: with the default `APIClient`, they are added only when the control plane advertises the required endpoints (unless `SkipEndpointChecks` is set); other client implementations register them unconditionally. The Insight @@ -118,6 +118,22 @@ Expose the ingested granular insight series (performance/test metrics parsed fro - `get_insight_metric_series` - Query a granular insight metric as a time series (values and trends over time) - `list_insight_executions` - List the workflow executions that produced a given insight metric +#### Insights Board Tools (9 tools) + +Manage Insights boards: saved, organization-wide dashboards made of reports (charts). Boards are organization-scoped, not environment-scoped; a report narrows itself to environments through its own `environment` filter. The Control Plane serves boards only to signed-in users, so these tools require a user session (`testkube login`, or the hosted endpoint with a user account) and return an actionable error for API tokens — the `APIClient` refuses a `tkcapi_` token before sending anything. They register unconditionally: the board routes do not answer HEAD, so they cannot be probed. + +- `list_boards` - List the boards visible to you (shared and private), with name, visibility and pinned filters +- `get_board` - Get a board with its reports in dashboard order and its layout +- `create_board` - Create an empty board (shared by default) +- `update_board` - Change a board's name, description, slug, visibility or layout +- `add_board_report` - Add a `pass-fail`, `executions`, `workflows` or `time-series` report +- `update_board_report` - Change a report's name, description, kind or params (merged by default) +- `remove_board_report` - Remove a report (destructive) +- `delete_board` - Delete a board and its reports (destructive) +- `render_board` - Run each report's query and return its numbers; `scope=environment` limits every report to the current environment, and `timeZone` (IANA, default UTC) anchors relative ranges at that zone's midnight, as the dashboard anchors them at the viewer's local midnight. + +Report params have no backend schema: the Control Plane stores them opaquely and only the dashboard interprets them. `pkg/mcp/boards` holds the port of the dashboard's rules — param validation and defaults, the report-to-query translation, and the layout — and both the `APIClient` and the Control Plane's `HandlerClient` must use it rather than reimplement it. The translation is pinned by `pkg/mcp/boards/testdata/translation_cases.json`. Two Control Plane behaviors the tools work around: older Control Planes clear the description of any update that omits it (current ones keep it), so each write reads the board and resends it; and deleting a report leaves its layout cell behind, so `remove_board_report` sends the recomputed layout with the delete. Because those values come from a read, every update also sends `expectedVersion`, the board's version as read: the Control Plane refuses the write with 409 if the board changed in between, and the tool reads the board again and rebuilds the write from it (up to three attempts) instead of overwriting the newer edit. `delete_board` resends nothing it read, so it is not conditional: it deletes by the ID it resolved, and the Control Plane checks delete rights against the board as it is then. + **Note for maintainers:** When adding new tools to `pkg/mcp/tools/`, ensure that: 1. The tool follows the interface-based design pattern (see existing tools for examples) diff --git a/pkg/mcp/api.go b/pkg/mcp/api.go index 7976dad54cd..6ee632a4bd2 100644 --- a/pkg/mcp/api.go +++ b/pkg/mcp/api.go @@ -4,6 +4,7 @@ import ( "bytes" "context" "encoding/json" + "errors" "fmt" "io" "net/http" @@ -14,6 +15,7 @@ import ( "time" "github.com/kubeshop/testkube/pkg/api/v1/testkube" + "github.com/kubeshop/testkube/pkg/mcp/boards" "github.com/kubeshop/testkube/pkg/mcp/tools" ) @@ -182,10 +184,7 @@ func (c *APIClient) makeRequest(ctx context.Context, apiReq APIRequest) (string, debugInfo.Data["responseBody"] = string(errBody) } } - if detail := extractErrorDetail(errBody); detail != "" { - return "", fmt.Errorf("API returned status %d: %s", resp.StatusCode, detail) - } - return "", fmt.Errorf("API returned status %d", resp.StatusCode) + return "", &apiStatusError{status: resp.StatusCode, detail: extractErrorDetail(errBody)} } // Read response body as string @@ -204,6 +203,21 @@ func (c *APIClient) makeRequest(ctx context.Context, apiReq APIRequest) (string, return string(bodyBytes), nil } +// apiStatusError is a non-2xx response. Its message is what callers have +// always seen; the status lets a method recognize one outcome, such as a 409 +// from a conditional board update, without matching on the text. +type apiStatusError struct { + status int + detail string +} + +func (e *apiStatusError) Error() string { + if e.detail != "" { + return fmt.Sprintf("API returned status %d: %s", e.status, e.detail) + } + return fmt.Sprintf("API returned status %d", e.status) +} + // maxErrorDetailLen bounds how much of an error response body is surfaced to the // caller, so a large/unexpected body can't blow up the tool result. const maxErrorDetailLen = 2048 @@ -1121,3 +1135,144 @@ func (c *APIClient) ListInsightExecutions(ctx context.Context, params tools.Insi QueryParams: queryParams, }) } + +// Insights board methods +// +// Boards are organization-scoped: they live under /organizations/{id}/boards, +// with no environment in the path or query. The Control Plane refuses API +// tokens on every board endpoint, and older versions answer them with a bare +// 500, so these methods refuse an API token themselves before sending anything. + +// apiTokenPrefix marks a Testkube API token, as opposed to a user access token. +const apiTokenPrefix = "tkcapi_" + +func (c *APIClient) requireUserSession() error { + if strings.HasPrefix(c.config.AccessToken, apiTokenPrefix) { + return tools.ErrBoardsRequireUser + } + return nil +} + +func (c *APIClient) ListBoards(ctx context.Context, params tools.ListBoardsParams) (string, error) { + if err := c.requireUserSession(); err != nil { + return "", err + } + queryParams := map[string]string{ + "name": params.Name, + "page": strconv.Itoa(params.Page), + "pageSize": strconv.Itoa(params.PageSize), + } + for key, set := range map[string]bool{ + "private": params.Private, + "shared": params.Shared, + "userFavorite": params.UserFavorite, + "orgFavorite": params.OrgFavorite, + } { + if set { + queryParams[key] = "true" + } + } + + return c.makeRequest(ctx, APIRequest{ + Method: http.MethodGet, + Path: "/boards", + Scope: ApiScopeOrg, + QueryParams: queryParams, + }) +} + +func (c *APIClient) GetBoard(ctx context.Context, board string) (string, error) { + if err := c.requireUserSession(); err != nil { + return "", err + } + return c.makeRequest(ctx, APIRequest{ + Method: http.MethodGet, + Path: "/boards/{boardID}", + Scope: ApiScopeOrg, + PathParams: map[string]string{"boardID": url.PathEscape(board)}, + }) +} + +func (c *APIClient) CheckBoardSlug(ctx context.Context, slug string) (bool, error) { + if err := c.requireUserSession(); err != nil { + return false, err + } + result, err := c.makeRequest(ctx, APIRequest{ + Method: http.MethodGet, + Path: "/board-slug", + Scope: ApiScopeOrg, + QueryParams: map[string]string{"slug": slug}, + }) + if err != nil { + return false, err + } + var availability struct { + Available bool `json:"available"` + } + if err := json.Unmarshal([]byte(result), &availability); err != nil { + return false, fmt.Errorf("failed to parse slug availability: %w", err) + } + return availability.Available, nil +} + +func (c *APIClient) CreateBoard(ctx context.Context, params tools.CreateBoardParams) (string, error) { + if err := c.requireUserSession(); err != nil { + return "", err + } + return c.makeRequest(ctx, APIRequest{ + Method: http.MethodPost, + Path: "/boards", + Scope: ApiScopeOrg, + Body: params, + }) +} + +func (c *APIClient) UpdateBoard(ctx context.Context, board string, request tools.UpdateBoardRequest) (string, error) { + if err := c.requireUserSession(); err != nil { + return "", err + } + result, err := c.makeRequest(ctx, APIRequest{ + Method: http.MethodPatch, + Path: "/boards/{boardID}", + Scope: ApiScopeOrg, + PathParams: map[string]string{"boardID": url.PathEscape(board)}, + Body: request, + }) + var statusErr *apiStatusError + if errors.As(err, &statusErr) && statusErr.status == http.StatusConflict { + return "", fmt.Errorf("%w: %v", tools.ErrBoardChanged, err) + } + return result, err +} + +func (c *APIClient) DeleteBoard(ctx context.Context, board string) error { + if err := c.requireUserSession(); err != nil { + return err + } + _, err := c.makeRequest(ctx, APIRequest{ + Method: http.MethodDelete, + Path: "/boards/{boardID}", + Scope: ApiScopeOrg, + PathParams: map[string]string{"boardID": url.PathEscape(board)}, + }) + return err +} + +// QueryBoardInsights runs the org-scoped insight query that renders a board +// report. Unlike the other insight methods it does not add the session's +// environment: a report carries its own environment filter, unless the query +// asks for the current environment. The insight endpoints themselves accept +// API tokens, but this is a board method and refuses them like the rest, as +// the Control Plane's client does. +func (c *APIClient) QueryBoardInsights(ctx context.Context, query boards.InsightQuery) (string, error) { + if err := c.requireUserSession(); err != nil { + return "", err + } + query = query.WithEnvironment(c.config.EnvId) + return c.makeRequest(ctx, APIRequest{ + Method: http.MethodGet, + Path: string(query.Endpoint), + Scope: ApiScopeOrg, + QueryParams: query.QueryParams(), + }) +} diff --git a/pkg/mcp/api_boards_test.go b/pkg/mcp/api_boards_test.go new file mode 100644 index 00000000000..4b689c89e93 --- /dev/null +++ b/pkg/mcp/api_boards_test.go @@ -0,0 +1,155 @@ +package mcp + +import ( + "context" + "encoding/json" + "errors" + "io" + "net/http" + "net/http/httptest" + "sync/atomic" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/kubeshop/testkube/pkg/mcp/boards" + "github.com/kubeshop/testkube/pkg/mcp/tools" +) + +func TestAPIClient_Boards_RequestShapes(t *testing.T) { + const ( + orgID = "org-1" + envID = "env-1" + ) + type seen struct { + method, path string + query map[string]string + body map[string]any + } + var got seen + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + got = seen{method: r.Method, path: r.URL.Path, query: map[string]string{}} + for k := range r.URL.Query() { + got.query[k] = r.URL.Query().Get(k) + } + if data, _ := io.ReadAll(r.Body); len(data) > 0 { + require.NoError(t, json.Unmarshal(data, &got.body)) + } + if r.URL.Path == "/organizations/"+orgID+"/board-slug" { + _, _ = io.WriteString(w, `{"available":true}`) + return + } + if r.Method == http.MethodDelete { + w.WriteHeader(http.StatusNoContent) + return + } + _, _ = io.WriteString(w, `{}`) + })) + defer server.Close() + + client := NewAPIClient(&MCPServerConfig{ControlPlaneUrl: server.URL, AccessToken: "user-token", OrgId: orgID, EnvId: envID}, server.Client()) + ctx := context.Background() + base := "/organizations/" + orgID + + _, err := client.ListBoards(ctx, tools.ListBoardsParams{Name: "q", Shared: true, Page: 1, PageSize: 20}) + require.NoError(t, err) + assert.Equal(t, seen{method: http.MethodGet, path: base + "/boards", query: map[string]string{"name": "q", "shared": "true", "page": "1", "pageSize": "20"}}, got) + + _, err = client.GetBoard(ctx, "my board") + require.NoError(t, err) + assert.Equal(t, base+"/boards/my board", got.path) + + available, err := client.CheckBoardSlug(ctx, "quality") + require.NoError(t, err) + assert.True(t, available) + assert.Equal(t, "quality", got.query["slug"]) + + _, err = client.CreateBoard(ctx, tools.CreateBoardParams{Name: "Q", IsPrivate: true}) + require.NoError(t, err) + assert.Equal(t, http.MethodPost, got.method) + assert.Equal(t, map[string]any{"name": "Q", "isPrivate": true}, got.body) + + description := "kept" + _, err = client.UpdateBoard(ctx, "tkcbrd_1", tools.UpdateBoardRequest{ + Description: &description, + Content: &tools.BoardContentPatch{Action: "create", ContentKind: "report", ContentData: &boards.ReportDraft{ + Kind: "workflows", Name: "W", Params: map[string]any{"duration": "month"}, + }}, + }) + require.NoError(t, err) + assert.Equal(t, http.MethodPatch, got.method) + assert.Equal(t, base+"/boards/tkcbrd_1", got.path) + assert.Equal(t, map[string]any{ + "description": "kept", + "content": map[string]any{ + "action": "create", "content_kind": "report", + // The description is always sent, so an empty one clears a report's. + "content_data": map[string]any{"kind": "workflows", "name": "W", "description": "", "params": map[string]any{"duration": "month"}}, + }, + }, got.body, "no layout key is sent when the layout is not changed") + + require.NoError(t, client.DeleteBoard(ctx, "tkcbrd_1")) + assert.Equal(t, http.MethodDelete, got.method) + + start := time.Date(2026, 9, 1, 0, 0, 0, 0, time.UTC) + query := boards.InsightQuery{Endpoint: boards.EndpointWorkflows, StartDate: start, EndDate: start.AddDate(0, 0, 7)} + _, err = client.QueryBoardInsights(ctx, query) + require.NoError(t, err) + assert.Equal(t, base+"/insights/workflows", got.path) + assert.NotContains(t, got.query, "env", "a report without an environment filter covers the whole organization") + assert.Equal(t, "2026-09-01T00:00:00Z", got.query["startDate"]) + + query.CurrentEnvironment = true + _, err = client.QueryBoardInsights(ctx, query) + require.NoError(t, err) + assert.Equal(t, envID, got.query["env"]) +} + +func TestAPIClient_Boards_RefuseAPITokenWithoutRequest(t *testing.T) { + var requests atomic.Int32 + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + requests.Add(1) + })) + defer server.Close() + + client := NewAPIClient(&MCPServerConfig{ControlPlaneUrl: server.URL, AccessToken: "tkcapi_abc", OrgId: "o", EnvId: "e"}, server.Client()) + ctx := context.Background() + + _, err := client.ListBoards(ctx, tools.ListBoardsParams{}) + assert.True(t, errors.Is(err, tools.ErrBoardsRequireUser)) + _, err = client.GetBoard(ctx, "b") + assert.True(t, errors.Is(err, tools.ErrBoardsRequireUser)) + _, err = client.CheckBoardSlug(ctx, "s") + assert.True(t, errors.Is(err, tools.ErrBoardsRequireUser)) + _, err = client.CreateBoard(ctx, tools.CreateBoardParams{Name: "n"}) + assert.True(t, errors.Is(err, tools.ErrBoardsRequireUser)) + _, err = client.UpdateBoard(ctx, "b", tools.UpdateBoardRequest{}) + assert.True(t, errors.Is(err, tools.ErrBoardsRequireUser)) + assert.True(t, errors.Is(client.DeleteBoard(ctx, "b"), tools.ErrBoardsRequireUser)) + _, err = client.QueryBoardInsights(ctx, boards.InsightQuery{Endpoint: boards.EndpointWorkflows}) + assert.True(t, errors.Is(err, tools.ErrBoardsRequireUser)) + assert.Zero(t, requests.Load(), "an API token must be refused before any request is sent") +} + +func TestAPIClient_UpdateBoard_ConflictIsErrBoardChanged(t *testing.T) { + var body map[string]any + server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { + data, _ := io.ReadAll(r.Body) + require.NoError(t, json.Unmarshal(data, &body)) + w.Header().Set("Content-Type", "application/problem+json") + w.WriteHeader(http.StatusConflict) + _, _ = io.WriteString(w, `{"title":"board changed","status":409,"detail":"the board changed after expectedVersion"}`) + })) + defer server.Close() + + client := NewAPIClient(&MCPServerConfig{ControlPlaneUrl: server.URL, AccessToken: "user-token", OrgId: "o", EnvId: "e"}, server.Client()) + name, version := "Renamed", int64(7) + _, err := client.UpdateBoard(context.Background(), "tkcbrd_1", tools.UpdateBoardRequest{Name: &name, ExpectedVersion: &version}) + + require.Error(t, err) + assert.True(t, errors.Is(err, tools.ErrBoardChanged)) + assert.Contains(t, err.Error(), "status 409", "the Control Plane's message is kept") + assert.Equal(t, float64(7), body["expectedVersion"]) +} diff --git a/pkg/mcp/boards/board.go b/pkg/mcp/boards/board.go new file mode 100644 index 00000000000..5b84bebae2b --- /dev/null +++ b/pkg/mcp/boards/board.go @@ -0,0 +1,77 @@ +package boards + +import ( + "encoding/json" + "fmt" + "strings" +) + +// Board is a board as the Control Plane returns it from get, create and +// update. +type Board struct { + ID string `json:"id"` + Slug string `json:"slug"` + Name string `json:"name"` + Description string `json:"description,omitempty"` + CreatedAt string `json:"createdAt,omitempty"` + UpdatedAt string `json:"updatedAt,omitempty"` + // Version is incremented by every change to the board. Nil when the + // Control Plane predates it. + Version *int64 `json:"version,omitempty"` + Creator string `json:"creator,omitempty"` + Shared bool `json:"shared"` + IsUserFavorite bool `json:"isUserFavorite,omitempty"` + IsOrgFavorite bool `json:"isOrgFavorite,omitempty"` + Layout json.RawMessage `json:"layout,omitempty"` + Content struct { + Reports []Report `json:"reports"` + } `json:"content"` +} + +// ParseBoard decodes a board response. +func ParseBoard(raw string) (*Board, error) { + if strings.TrimSpace(raw) == "" { + return nil, fmt.Errorf("empty board response") + } + var b Board + if err := json.Unmarshal([]byte(raw), &b); err != nil { + return nil, fmt.Errorf("failed to parse board: %w", err) + } + return &b, nil +} + +// ReportIDs returns the IDs of the board's reports in storage order. +func (b *Board) ReportIDs() []string { + ids := make([]string, 0, len(b.Content.Reports)) + for _, r := range b.Content.Reports { + ids = append(ids, r.ID) + } + return ids +} + +// FindReport returns the board's report with the given ID. +func (b *Board) FindReport(id string) (*Report, bool) { + for i := range b.Content.Reports { + if b.Content.Reports[i].ID == id { + return &b.Content.Reports[i], true + } + } + return nil, false +} + +// OrderedReports returns the board's reports in the order the dashboard shows +// them, and the IDs of the reports its layout does not place. +func (b *Board) OrderedReports() (ordered []Report, unplaced []string, err error) { + layout, err := ParseLayout(b.Layout) + if err != nil { + return nil, nil, err + } + placed, unplaced := layout.Order(b.ReportIDs()) + for _, ids := range [][]string{placed, unplaced} { + for _, id := range ids { + r, _ := b.FindReport(id) + ordered = append(ordered, *r) + } + } + return ordered, unplaced, nil +} diff --git a/pkg/mcp/boards/boards_test.go b/pkg/mcp/boards/boards_test.go new file mode 100644 index 00000000000..c3e50761a97 --- /dev/null +++ b/pkg/mcp/boards/boards_test.go @@ -0,0 +1,310 @@ +package boards + +import ( + "encoding/json" + "os" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// translationCase pins the query the dashboard issues for a report. The cases +// live in a JSON fixture so the dashboard's own tests can replay them. +type translationCase struct { + Name string `json:"name"` + Kind string `json:"kind"` + Now time.Time `json:"now"` + TimeZone string `json:"timeZone"` + Params map[string]any `json:"params"` + Endpoint string `json:"endpoint"` + Query map[string]string `json:"query"` +} + +func TestBuildQuery_TranslationCases(t *testing.T) { + data, err := os.ReadFile("testdata/translation_cases.json") + require.NoError(t, err) + var cases []translationCase + require.NoError(t, json.Unmarshal(data, &cases)) + require.NotEmpty(t, cases) + + for _, tc := range cases { + t.Run(tc.Name, func(t *testing.T) { + loc, err := ParseTimeZone(tc.TimeZone) + require.NoError(t, err) + q, err := BuildQuery(tc.Kind, tc.Params, QueryOptions{Now: tc.Now, Location: loc}) + require.NoError(t, err) + assert.Equal(t, tc.Endpoint, string(q.Endpoint)) + assert.Equal(t, tc.Query, q.QueryParams()) + }) + } +} + +func TestBuildQuery_CurrentEnvironment(t *testing.T) { + params := map[string]any{ + "filter": []any{map[string]any{"filterConfigurationKey": "environment", "id": "1", "operator": "is", "value": []any{"tkcenv_other"}}}, + } + + q, err := BuildQuery(KindWorkflows, params, QueryOptions{CurrentEnvironment: true}) + require.NoError(t, err) + assert.True(t, q.CurrentEnvironment) + assert.Empty(t, q.Env, "the report's own environment filter must not leak into a current-environment query") + assert.Equal(t, "tkcenv_mine", q.WithEnvironment("tkcenv_mine").QueryParams()["env"]) + + q, err = BuildQuery(KindWorkflows, params, QueryOptions{}) + require.NoError(t, err) + assert.Equal(t, "tkcenv_other", q.WithEnvironment("tkcenv_mine").Env, "WithEnvironment only applies to current-environment queries") +} + +func TestBuildQuery_Errors(t *testing.T) { + _, err := BuildQuery("pie", nil, QueryOptions{}) + assert.ErrorContains(t, err, "unknown report kind") + + _, err = BuildQuery(KindTimeSeries, map[string]any{}, QueryOptions{}) + assert.ErrorContains(t, err, "no measure") + + _, err = BuildQuery(KindWorkflows, map[string]any{"from": "yesterday"}, QueryOptions{}) + assert.ErrorContains(t, err, "RFC3339") +} + +func TestNormalizeReport_Defaults(t *testing.T) { + tests := []struct { + kind string + want map[string]any + }{ + {KindPassFail, map[string]any{"filter": []any{}, "duration": "month", "measure": "ratio"}}, + {KindExecutions, map[string]any{"filter": []any{}, "groupBy": "status", "measure": "count"}}, + {KindWorkflows, map[string]any{"filter": []any{}, "duration": "month"}}, + {KindTimeSeries, map[string]any{"filter": []any{}, "duration": "week", "measure": "execution-count", "aggregate": "sum", "segment": "status", "chartType": "bar"}}, + } + for _, tt := range tests { + t.Run(tt.kind, func(t *testing.T) { + got, err := NormalizeReport(tt.kind, nil) + require.NoError(t, err) + assert.Equal(t, tt.want, got) + }) + } +} + +func TestNormalizeReport_TimeSeriesSegment(t *testing.T) { + got, err := NormalizeReport(KindTimeSeries, map[string]any{"measure": "cpu-millicores-max"}) + require.NoError(t, err) + assert.NotContains(t, got, "segment", "only execution-count starts segmented by status") + + got, err = NormalizeReport(KindTimeSeries, map[string]any{"segment": ""}) + require.NoError(t, err) + assert.NotContains(t, got, "segment", "an explicit empty segment means no segment") +} + +func TestNormalizeReport_Validation(t *testing.T) { + tests := []struct { + name string + kind string + params map[string]any + want string + }{ + {"unknown kind", "pie", nil, "unknown report kind"}, + {"bad duration", KindWorkflows, map[string]any{"duration": "year"}, "param duration must be one of"}, + {"bad pass-fail measure", KindPassFail, map[string]any{"measure": "count"}, "param measure must be one of ratio"}, + {"bad executions measure", KindExecutions, map[string]any{"measure": "ratio"}, "param measure must be one of count"}, + {"empty groupBy", KindExecutions, map[string]any{"groupBy": " "}, "groupBy must be a non-empty string"}, + {"bad aggregate", KindTimeSeries, map[string]any{"aggregate": "median"}, "param aggregate must be one of"}, + {"bad chart type", KindTimeSeries, map[string]any{"chartType": "pie"}, "param chartType must be one of"}, + {"bad overlay", KindTimeSeries, map[string]any{"overlaySuccessRate": "yes"}, "overlaySuccessRate must be a boolean"}, + {"bad from", KindWorkflows, map[string]any{"from": "2026-01-01"}, "param from must be an RFC3339"}, + {"filter not a list", KindWorkflows, map[string]any{"filter": "workflow=a"}, "param filter must be an array"}, + {"filter without key", KindWorkflows, map[string]any{"filter": []any{map[string]any{"value": "a"}}}, "missing filterConfigurationKey"}, + {"filter with bad value", KindWorkflows, map[string]any{"filter": []any{map[string]any{"filterConfigurationKey": "workflow", "value": 3}}}, "invalid value"}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + _, err := NormalizeReport(tt.kind, tt.params) + assert.ErrorContains(t, err, tt.want) + }) + } +} + +func TestNormalizeReport_FillsFilterIDsAndOperators(t *testing.T) { + got, err := NormalizeReport(KindWorkflows, map[string]any{ + "filter": []any{map[string]any{"filterConfigurationKey": "workflow", "value": []any{"a"}}}, + }) + require.NoError(t, err) + filters, err := ParseFilters(got["filter"]) + require.NoError(t, err) + require.Len(t, filters, 1) + assert.NotEmpty(t, filters[0].ID) + assert.Equal(t, "is", filters[0].Operator) +} + +func TestNormalizeReport_PreservesUnknownKeys(t *testing.T) { + got, err := NormalizeReport(KindWorkflows, map[string]any{"futureKey": 1}) + require.NoError(t, err) + assert.Equal(t, 1, got["futureKey"], "keys written by a newer dashboard must survive an MCP edit") +} + +func TestCheckParamKeys(t *testing.T) { + assert.NoError(t, CheckParamKeys(KindExecutions, map[string]any{"groupBy": "workflow", "measure": "count", "duration": "day"})) + assert.ErrorContains(t, CheckParamKeys(KindWorkflows, map[string]any{"measure": "count", "grupBy": "x"}), "unknown params for a workflows report: grupBy, measure") +} + +func TestApplyFilters(t *testing.T) { + params := map[string]any{ + "filter": []any{ + map[string]any{"filterConfigurationKey": "workflow", "id": "w", "operator": "is", "value": []any{"old"}}, + map[string]any{"filterConfigurationKey": "status", "id": "s", "operator": "is", "value": []any{"failed"}}, + map[string]any{"filterConfigurationKey": "labels-v2", "id": "l", "operator": "is", "value": []any{"a=b"}}, + }, + } + require.NoError(t, ApplyFilters(params, map[string][]string{ + "workflow": {"new", " "}, + "status": {}, + "labels": {"team=core"}, + })) + + filters, err := ParseFilters(params["filter"]) + require.NoError(t, err) + assert.Equal(t, []string{"new"}, FilterValues(filters, FilterWorkflow)) + assert.Empty(t, FilterValues(filters, FilterStatus), "an empty list removes the key's filters") + assert.Equal(t, []string{"team=core"}, FilterValues(filters, FilterLabels), "labels is an alias of labels-v2 and replaces it") +} + +func TestLayout(t *testing.T) { + raw := json.RawMessage(`{"version":1,"rows":[{"id":"r1","cells":[{"id":"a"},{"id":"b"}]},{"id":"r2","cells":[{"id":"c"}]}]}`) + layout, err := ParseLayout(raw) + require.NoError(t, err) + + t.Run("without drops the cell and empty rows", func(t *testing.T) { + assert.Equal(t, [][]string{{"a", "b"}}, layout.Without("c").RowIDs()) + assert.Equal(t, [][]string{{"b"}, {"c"}}, layout.Without("a").RowIDs()) + }) + + t.Run("order puts unplaced reports last", func(t *testing.T) { + placed, unplaced := layout.Order([]string{"c", "d", "b", "a"}) + assert.Equal(t, []string{"a", "b", "c"}, placed) + assert.Equal(t, []string{"d"}, unplaced) + }) + + t.Run("validate", func(t *testing.T) { + ok := Layout{Version: 1, Rows: []Row{{Cells: []Cell{{ID: "b"}, {ID: "a"}}}}} + require.NoError(t, ok.Validate([]string{"a", "b"})) + assert.NotEmpty(t, ok.Rows[0].ID, "missing row IDs are generated") + + assert.ErrorContains(t, (&Layout{Version: 2}).Validate(nil), "version must be 1") + assert.ErrorContains(t, (&Layout{Version: 1, Rows: []Row{{Cells: []Cell{{ID: "x"}}}}}).Validate([]string{"a"}), "unknown report") + assert.ErrorContains(t, (&Layout{Version: 1, Rows: []Row{{Cells: []Cell{{ID: "a"}, {ID: "a"}}}}}).Validate([]string{"a"}), "more than once") + assert.ErrorContains(t, (&Layout{Version: 1, Rows: []Row{{Cells: []Cell{{ID: "a"}}}}}).Validate([]string{"a", "b"}), "leaves out reports b") + assert.ErrorContains(t, (&Layout{Version: 1, Rows: []Row{{}}}).Validate(nil), "has no cells") + }) + + t.Run("an empty layout parses as version 1", func(t *testing.T) { + for _, raw := range []string{"", "null", "{}"} { + l, err := ParseLayout(json.RawMessage(raw)) + require.NoError(t, err) + assert.Equal(t, LayoutVersion, l.Version) + assert.Empty(t, l.Rows) + } + }) +} + +func TestBoardOrderedReports(t *testing.T) { + b, err := ParseBoard(`{"id":"b1","name":"B","layout":{"version":1,"rows":[{"id":"r","cells":[{"id":"y"}]}]}, + "content":{"reports":[{"id":"x","kind":"workflows"},{"id":"y","kind":"pass-fail"}]}}`) + require.NoError(t, err) + + ordered, unplaced, err := b.OrderedReports() + require.NoError(t, err) + require.Len(t, ordered, 2) + assert.Equal(t, "y", ordered[0].ID) + assert.Equal(t, "x", ordered[1].ID) + assert.Equal(t, []string{"x"}, unplaced) +} + +func TestParseTimeZone(t *testing.T) { + loc, err := ParseTimeZone("") + require.NoError(t, err) + assert.Equal(t, time.UTC, loc) + + loc, err = ParseTimeZone(" Asia/Kolkata ") + require.NoError(t, err) + assert.Equal(t, "Asia/Kolkata", loc.String()) + + _, err = ParseTimeZone("Mars/Olympus") + assert.ErrorContains(t, err, "IANA time zone") +} + +func TestValidateReport_AddsNoDefaults(t *testing.T) { + // A stored report the dashboard renders from its missing params: an edit + // must not fill in creation defaults that render differently. + tests := []struct { + kind string + params map[string]any + absent []string + }{ + {KindPassFail, map[string]any{"measure": "failed-count"}, []string{"duration"}}, + {KindWorkflows, map[string]any{}, []string{"duration"}}, + {KindTimeSeries, map[string]any{"measure": DefaultTimeSeriesMeasure}, []string{"segment", "duration", "chartType", "aggregate"}}, + {KindExecutions, map[string]any{}, []string{"groupBy", "measure"}}, + } + for _, tt := range tests { + t.Run(tt.kind, func(t *testing.T) { + got, err := ValidateReport(tt.kind, tt.params) + require.NoError(t, err) + for _, key := range tt.absent { + assert.NotContains(t, got, key) + } + }) + } + + t.Run("the rendered window is the one the report had", func(t *testing.T) { + now := time.Date(2026, 9, 28, 12, 0, 0, 0, time.UTC) + stored := map[string]any{"measure": "ratio"} + before, err := BuildQuery(KindPassFail, stored, QueryOptions{Now: now}) + require.NoError(t, err) + edited, err := ValidateReport(KindPassFail, MergeParams(stored, map[string]any{"measure": "failed-count"})) + require.NoError(t, err) + after, err := BuildQuery(KindPassFail, edited, QueryOptions{Now: now}) + require.NoError(t, err) + assert.Equal(t, before.StartDate, after.StartDate, "an unrelated edit must not change the date window") + }) + + t.Run("it still refuses bad values", func(t *testing.T) { + _, err := ValidateReport(KindPassFail, map[string]any{"duration": "year"}) + assert.ErrorContains(t, err, "param duration must be one of") + _, err = ValidateReport(KindTimeSeries, map[string]any{"segment": 3}) + assert.ErrorContains(t, err, "param segment must be a string") + }) +} + +func TestParseFilters_LabelShapes(t *testing.T) { + label := func(operator string, value any) []any { + return []any{map[string]any{"filterConfigurationKey": FilterLabels, "id": "l", "operator": operator, "value": value}} + } + valid := [][]any{ + label("is", []any{"team=core"}), + label("where", map[string]any{"labelKey": "tier", "labelOperator": "contains", "labelValue": "gold"}), + label("where", map[string]any{"labelKey": "owner", "labelOperator": "exists"}), + } + for _, f := range valid { + _, err := ParseFilters(f) + assert.NoError(t, err) + } + invalid := map[string][]any{ + "a string with is": label("is", "team=core"), + "an array with where": label("where", []any{"team=core"}), + "where without a key": label("where", map[string]any{"labelOperator": "exists"}), + "where with a bad op": label("where", map[string]any{"labelKey": "tier", "labelOperator": "equals"}), + "an unsupported operator": label("contains", []any{"team=core"}), + } + for name, f := range invalid { + t.Run(name, func(t *testing.T) { + _, err := ParseFilters(f) + assert.ErrorContains(t, err, "labels-v2") + }) + } + + t.Run("a stored report with a bad label filter fails to render rather than dropping it", func(t *testing.T) { + _, err := BuildQuery(KindWorkflows, map[string]any{"filter": label("is", "team=core")}, QueryOptions{}) + assert.ErrorContains(t, err, "invalid value") + }) +} diff --git a/pkg/mcp/boards/layout.go b/pkg/mcp/boards/layout.go new file mode 100644 index 00000000000..7e94bfa1a43 --- /dev/null +++ b/pkg/mcp/boards/layout.go @@ -0,0 +1,142 @@ +package boards + +import ( + "encoding/json" + "fmt" + "strings" +) + +// LayoutVersion is the only board layout version the dashboard understands. +const LayoutVersion = 1 + +// Layout places a board's reports on rows. Each cell holds one report ID. +type Layout struct { + Version int `json:"version"` + Rows []Row `json:"rows"` +} + +// Row is one row of a board layout. +type Row struct { + ID string `json:"id"` + Cells []Cell `json:"cells"` +} + +// Cell is one slot of a row, holding a report. +type Cell struct { + ID string `json:"id"` +} + +// ParseLayout decodes a board layout. A missing or null layout is an empty +// version 1 layout, as the Control Plane creates for a new board. +func ParseLayout(raw json.RawMessage) (Layout, error) { + trimmed := strings.TrimSpace(string(raw)) + if trimmed == "" || trimmed == "null" || trimmed == "{}" { + return Layout{Version: LayoutVersion, Rows: []Row{}}, nil + } + var l Layout + if err := json.Unmarshal(raw, &l); err != nil { + return Layout{}, fmt.Errorf("layout is malformed: %w", err) + } + if l.Rows == nil { + l.Rows = []Row{} + } + return l, nil +} + +// Without returns the layout with the report's cell removed and any row left +// empty dropped. The Control Plane does not touch the layout when a report is +// deleted, so the caller must send this alongside the delete. +func (l Layout) Without(reportID string) Layout { + out := Layout{Version: l.Version, Rows: make([]Row, 0, len(l.Rows))} + for _, row := range l.Rows { + cells := make([]Cell, 0, len(row.Cells)) + for _, c := range row.Cells { + if c.ID != reportID { + cells = append(cells, c) + } + } + if len(cells) > 0 { + out.Rows = append(out.Rows, Row{ID: row.ID, Cells: cells}) + } + } + return out +} + +// Validate checks a caller-supplied layout against the board's reports: it +// must be version 1, reference only existing reports, place each at most once, +// and leave none out. Missing row IDs are generated. +func (l *Layout) Validate(reportIDs []string) error { + if l.Version != LayoutVersion { + return fmt.Errorf("layout version must be %d, got %d", LayoutVersion, l.Version) + } + known := map[string]bool{} + for _, id := range reportIDs { + known[id] = true + } + seen := map[string]bool{} + for i := range l.Rows { + if l.Rows[i].ID == "" { + l.Rows[i].ID = newID() + } + if len(l.Rows[i].Cells) == 0 { + return fmt.Errorf("layout row %d has no cells; drop empty rows", i) + } + for _, c := range l.Rows[i].Cells { + if !known[c.ID] { + return fmt.Errorf("layout references unknown report %q", c.ID) + } + if seen[c.ID] { + return fmt.Errorf("layout places report %q more than once", c.ID) + } + seen[c.ID] = true + } + } + var missing []string + for _, id := range reportIDs { + if !seen[id] { + missing = append(missing, id) + } + } + if len(missing) > 0 { + return fmt.Errorf("layout leaves out reports %s; every report must be placed (use remove_board_report to delete one)", strings.Join(missing, ", ")) + } + return nil +} + +// Order returns the report IDs in the order the dashboard shows them - row by +// row, cell by cell - followed by the reports the layout does not place, which +// the dashboard does not show at all. +func (l Layout) Order(reportIDs []string) (placed, unplaced []string) { + known := map[string]bool{} + for _, id := range reportIDs { + known[id] = true + } + seen := map[string]bool{} + for _, row := range l.Rows { + for _, c := range row.Cells { + if known[c.ID] && !seen[c.ID] { + placed = append(placed, c.ID) + seen[c.ID] = true + } + } + } + for _, id := range reportIDs { + if !seen[id] { + unplaced = append(unplaced, id) + } + } + return placed, unplaced +} + +// RowIDs returns the report IDs of each row, for a compact view of the layout. +func (l Layout) RowIDs() [][]string { + out := make([][]string, 0, len(l.Rows)) + for _, row := range l.Rows { + ids := make([]string, 0, len(row.Cells)) + for _, c := range row.Cells { + ids = append(ids, c.ID) + } + out = append(out, ids) + } + return out +} diff --git a/pkg/mcp/boards/query.go b/pkg/mcp/boards/query.go new file mode 100644 index 00000000000..7bfb569b643 --- /dev/null +++ b/pkg/mcp/boards/query.go @@ -0,0 +1,249 @@ +package boards + +import ( + "encoding/json" + "fmt" + "strings" + "time" +) + +// Endpoint is the org-scoped insight endpoint that renders a report kind, +// relative to /organizations/{id}. +type Endpoint string + +const ( + EndpointStats Endpoint = "/insights/stats" + EndpointExecutions Endpoint = "/insights/executions" + EndpointWorkflows Endpoint = "/insights/workflows" + EndpointSeries Endpoint = "/insights/series" +) + +// durationMinutes mirrors the dashboard's DURATION_MAP. A quarter is 4 x 30 +// days there, not 90, and the port keeps it that way so both sides agree. +var durationMinutes = map[string]int{ + "day": 24 * 60, + "week": 24 * 60 * 7, + "month": 24 * 60 * 30, + "quarter": 24 * 60 * 30 * 4, +} + +// InsightQuery is the request that renders one report. Empty fields are not +// sent. Env comes from the report's environment filter; a report without one +// covers the whole organization, which is what the dashboard shows for it. +type InsightQuery struct { + Endpoint Endpoint + StartDate time.Time + EndDate time.Time + + Env string + Workflow string + Status string + Selector string + TagFilter string + + // executions + GroupBy string + + // time-series + Measure string + Aggregate string + Segment string + IdentityFilters string + + // CurrentEnvironment asks the client to scope the query to the environment + // the MCP session runs in, replacing the report's own environment filter. + // Only the client knows that environment, so it applies it with + // WithEnvironment before sending the query. + CurrentEnvironment bool +} + +// WithEnvironment returns the query scoped to env when it asks for the current +// environment, and unchanged otherwise. +func (q InsightQuery) WithEnvironment(env string) InsightQuery { + if q.CurrentEnvironment { + q.Env = env + } + return q +} + +// QueryParams encodes the query as URL query parameters. Empty values are +// omitted. +func (q InsightQuery) QueryParams() map[string]string { + params := map[string]string{ + "startDate": q.StartDate.UTC().Format(time.RFC3339), + "endDate": q.EndDate.UTC().Format(time.RFC3339), + } + set := func(key, value string) { + if value != "" { + params[key] = value + } + } + set("env", q.Env) + set("workflow", q.Workflow) + set("status", q.Status) + set("selector", q.Selector) + set("tagFilter", q.TagFilter) + set("groupBy", q.GroupBy) + set("measure", q.Measure) + set("aggregate", q.Aggregate) + set("segment", q.Segment) + set("identityFilters", q.IdentityFilters) + return params +} + +// QueryOptions controls how a report is turned into a query. +type QueryOptions struct { + // Now anchors relative durations. Zero means time.Now(). + Now time.Time + // CurrentEnvironment replaces the report's own environment filter with the + // MCP session's environment (see InsightQuery.CurrentEnvironment). + CurrentEnvironment bool + // Location is the time zone relative durations are anchored in. The + // dashboard anchors them to the viewer's local midnight, so this must be + // the viewer's zone for the numbers to match. Nil means UTC. + Location *time.Location +} + +// BuildQuery translates a report into the insight query the dashboard issues +// to render it (a port of mapBaseFilterToQueryParams, useInsightsBaseFilter and +// useTimeSeriesBaseFilter). +func BuildQuery(kind string, params map[string]any, opts QueryOptions) (InsightQuery, error) { + if !IsKind(kind) { + return InsightQuery{}, unknownKindError(kind) + } + now := opts.Now + if now.IsZero() { + now = time.Now() + } + + filters, err := ParseFilters(params["filter"]) + if err != nil { + return InsightQuery{}, err + } + start, end, err := DateRange(stringParam(params, "duration"), stringParam(params, "from"), stringParam(params, "to"), now, opts.Location) + if err != nil { + return InsightQuery{}, err + } + + q := InsightQuery{StartDate: start, EndDate: end} + q.Env = strings.Join(FilterValues(filters, FilterEnvironment), ",") + if opts.CurrentEnvironment { + q.Env = "" + q.CurrentEnvironment = true + } + q.Workflow = strings.Join(compact(FilterValues(filters, FilterWorkflow)), ",") + q.Status = strings.Join(FilterValues(filters, FilterStatus), ",") + q.Selector = strings.Join(compact(FilterValues(filters, FilterLabels)), ",") + // Tag values may contain commas, so tags travel as a JSON array. + if tags := compact(FilterValues(filters, FilterTags)); len(tags) > 0 { + data, _ := json.Marshal(tags) + q.TagFilter = string(data) + } + + switch kind { + case KindPassFail: + q.Endpoint = EndpointStats + // The stats endpoint has no status filter; the dashboard sends one and + // the backend ignores it. + q.Status = "" + case KindExecutions: + q.Endpoint = EndpointExecutions + q.GroupBy = stringParam(params, "groupBy") + if q.GroupBy == "" { + q.GroupBy = "status" + } + case KindWorkflows: + q.Endpoint = EndpointWorkflows + case KindTimeSeries: + q.Endpoint = EndpointSeries + q.Measure = stringParam(params, "measure") + if q.Measure == "" { + // The dashboard does not query a time-series report until a + // measure is chosen. + return InsightQuery{}, fmt.Errorf("time-series report has no measure; set params.measure") + } + q.Aggregate = stringParam(params, "aggregate") + if q.Aggregate == "" { + q.Aggregate = "sum" + } + q.Segment = stringParam(params, "segment") + if identity := identityFilters(filters); len(identity) > 0 { + data, _ := json.Marshal(identity) + q.IdentityFilters = string(data) + } + } + return q, nil +} + +// DateRange resolves a report's duration/from/to into the queried range, as +// the dashboard's getDateParams does: +// - from and to: used as given; +// - only from: [from, from + duration]; +// - only to: [to - duration, to]; +// - neither: the duration ending at the start of tomorrow. +// +// A missing or "custom" duration means a week. "Tomorrow" is taken in loc, +// as the dashboard takes it in the browser's time zone; nil means UTC. The +// duration is then subtracted as a fixed number of minutes, as the dashboard +// does, so a range that crosses a daylight-saving change starts an hour off +// local midnight there too. +func DateRange(duration, from, to string, now time.Time, loc *time.Location) (time.Time, time.Time, error) { + var fromT, toT time.Time + var err error + if from != "" { + if fromT, err = time.Parse(time.RFC3339, from); err != nil { + return time.Time{}, time.Time{}, fmt.Errorf("param from must be an RFC3339 timestamp: %w", err) + } + } + if to != "" { + if toT, err = time.Parse(time.RFC3339, to); err != nil { + return time.Time{}, time.Time{}, fmt.Errorf("param to must be an RFC3339 timestamp: %w", err) + } + } + if from != "" && to != "" { + return fromT, toT, nil + } + + minutes, ok := durationMinutes[duration] + if !ok { + minutes = durationMinutes["week"] + } + d := time.Duration(minutes) * time.Minute + switch { + case from != "": + return fromT, fromT.Add(d), nil + case to != "": + return toT.Add(-d), toT, nil + } + if loc == nil { + loc = time.UTC + } + y, m, day := now.In(loc).Date() + eod := time.Date(y, m, day+1, 0, 0, 0, 0, loc) + return eod.Add(-d), eod, nil +} + +// identityFilters collects the non-base filter keys of a time-series report, +// a port of getTimeSeriesIdentityFilters. +func identityFilters(filters []Filter) map[string][]string { + base := map[string]bool{FilterEnvironment: true, FilterWorkflow: true, FilterStatus: true, FilterLabels: true, FilterTags: true} + result := map[string][]string{} + for _, f := range filters { + key := f.FilterConfigurationKey + if base[key] { + continue + } + if _, done := result[key]; done { + continue + } + if values := compact(FilterValues(filters, key)); len(values) > 0 { + result[key] = values + } + } + return result +} + +func stringParam(params map[string]any, key string) string { + s, _ := params[key].(string) + return s +} diff --git a/pkg/mcp/boards/reports.go b/pkg/mcp/boards/reports.go new file mode 100644 index 00000000000..6b73a6e2f4d --- /dev/null +++ b/pkg/mcp/boards/reports.go @@ -0,0 +1,530 @@ +// Package boards holds the transport-free logic behind the Insights board MCP +// tools: the shape of a report's params, the translation of a report into the +// insight query that renders it, and the board layout. +// +// Report params have no backend schema - the Control Plane stores them as an +// opaque JSON object and only the dashboard interprets them. The rules here are +// a port of the dashboard's report types and filter mapping, so a report the +// MCP writes renders and edits in the UI, and render_board shows the numbers +// the UI shows. Both the CLI client and the Control Plane's hosted MCP use this +// package; neither side should reimplement it. +package boards + +import ( + "crypto/rand" + "encoding/hex" + "encoding/json" + "fmt" + "slices" + "sort" + "strings" + "time" +) + +// Report kinds, matching the dashboard's report registry. +const ( + KindTimeSeries = "time-series" + KindExecutions = "executions" + KindPassFail = "pass-fail" + KindWorkflows = "workflows" +) + +// Kinds lists every report kind the dashboard can render. +var Kinds = []string{KindTimeSeries, KindExecutions, KindPassFail, KindWorkflows} + +// Filter configuration keys the dashboard treats as base filters. Any other key +// on a time-series report is an identity filter on the granular series. +const ( + FilterEnvironment = "environment" + FilterWorkflow = "workflow" + FilterStatus = "status" + FilterLabels = "labels-v2" + FilterTags = "tags" +) + +var ( + durations = []string{"day", "week", "month", "quarter"} + passFailMeasures = []string{"ratio", "failed-count", "total-count"} + executionsMeasures = []string{"count", "duration"} + // The dashboard offers "last" as well; the backend falls back to sum for it. + timeSeriesAggregates = []string{"sum", "avg", "min", "max", "count", "last"} + timeSeriesChartTypes = []string{"bar", "bar-grouped", "bar-normalized", "line-stacked", "line", "area-normalized", "heatmap", "horizon"} +) + +// DefaultTimeSeriesMeasure is the measure a new time-series report starts with. +const DefaultTimeSeriesMeasure = "execution-count" + +// paramKeys lists the params each kind understands. +var paramKeys = map[string][]string{ + KindPassFail: {"filter", "duration", "from", "to", "measure"}, + KindExecutions: {"filter", "duration", "from", "to", "groupBy", "measure"}, + KindWorkflows: {"filter", "duration", "from", "to"}, + KindTimeSeries: {"filter", "duration", "from", "to", "measure", "aggregate", "segment", "chartType", "overlaySuccessRate"}, +} + +// Report is one report ("analysis") on a board, as the Control Plane returns it. +type Report struct { + ID string `json:"id"` + Kind string `json:"kind"` + Name string `json:"name,omitempty"` + Description string `json:"description,omitempty"` + Params map[string]any `json:"params,omitempty"` +} + +// ReportDraft is the content of a report create or update request. +type ReportDraft struct { + Kind string `json:"kind"` + Name string `json:"name"` + Description string `json:"description"` + Params map[string]any `json:"params"` +} + +// Filter is one entry of a report's params.filter, in the dashboard's format. +type Filter struct { + FilterConfigurationKey string `json:"filterConfigurationKey"` + ID string `json:"id"` + Operator string `json:"operator"` + Value json.RawMessage `json:"value"` +} + +// LabelFilterValue is the value of an advanced ("where") label filter. +type LabelFilterValue struct { + LabelKey string `json:"labelKey"` + LabelOperator string `json:"labelOperator"` + LabelValue string `json:"labelValue"` +} + +// IsKind reports whether kind is a report kind the dashboard can render. +func IsKind(kind string) bool { + return slices.Contains(Kinds, kind) +} + +// CheckParamKeys rejects params the kind does not understand, so a typo in a +// caller's input is reported instead of being stored and silently ignored. +// It is applied to caller input only: a report written by a newer dashboard may +// carry keys this package does not know, and those are preserved. +func CheckParamKeys(kind string, params map[string]any) error { + allowed, ok := paramKeys[kind] + if !ok { + return unknownKindError(kind) + } + var unknown []string + for key := range params { + if !slices.Contains(allowed, key) { + unknown = append(unknown, key) + } + } + if len(unknown) > 0 { + sort.Strings(unknown) + return fmt.Errorf("unknown params for a %s report: %s (allowed: %s)", kind, strings.Join(unknown, ", "), strings.Join(allowed, ", ")) + } + return nil +} + +// FiltersToParams converts the convenience filters object ({"workflow": [...], +// "labels": [...], ...}) into dashboard filters. "labels" is accepted as an +// alias of the dashboard's "labels-v2" key. Keys are processed in sorted order +// so the output is deterministic. +func FiltersToParams(filters map[string][]string) []Filter { + keys := make([]string, 0, len(filters)) + for key := range filters { + keys = append(keys, key) + } + sort.Strings(keys) + + result := make([]Filter, 0, len(keys)) + for _, key := range keys { + values := compact(filters[key]) + if len(values) == 0 { + continue + } + value, _ := json.Marshal(values) + result = append(result, Filter{ + FilterConfigurationKey: canonicalFilterKey(key), + ID: newID(), + Operator: "is", + Value: value, + }) + } + return result +} + +// ApplyFilters merges convenience filters into params.filter: every key named +// in filters replaces the existing entries for that key, and an empty list +// removes them. Keys not named are left untouched. +func ApplyFilters(params map[string]any, filters map[string][]string) error { + if len(filters) == 0 { + return nil + } + existing, err := ParseFilters(params["filter"]) + if err != nil { + return err + } + replaced := map[string]bool{} + for key := range filters { + replaced[canonicalFilterKey(key)] = true + } + merged := make([]Filter, 0, len(existing)+len(filters)) + for _, f := range existing { + if !replaced[f.FilterConfigurationKey] { + merged = append(merged, f) + } + } + merged = append(merged, FiltersToParams(filters)...) + params["filter"] = filtersToAny(merged) + return nil +} + +// NormalizeReport validates the params of a new report and fills in the +// defaults the dashboard gives a new report of that kind, returning a new +// params map. Unknown keys are preserved (see CheckParamKeys). +// +// It is for creating a report, or replacing its params wholesale. To edit an +// existing report use ValidateReport: some creation defaults differ from how +// the dashboard renders a param that is missing (a pass-fail report without a +// duration renders a week but is created with a month), so applying them to +// a stored report would change what it shows. +func NormalizeReport(kind string, params map[string]any) (map[string]any, error) { + if !IsKind(kind) { + return nil, unknownKindError(kind) + } + out := copyParams(params) + switch kind { + case KindPassFail: + setDefault(out, "duration", "month") + setDefault(out, "measure", "ratio") + case KindExecutions: + setDefault(out, "groupBy", "status") + setDefault(out, "measure", "count") + case KindWorkflows: + setDefault(out, "duration", "month") + case KindTimeSeries: + setDefault(out, "duration", "week") + setDefault(out, "measure", DefaultTimeSeriesMeasure) + setDefault(out, "aggregate", "sum") + // A new report is segmented by status only for the measure whose + // segments are execution statuses; an explicit "" means "no segment". + if _, ok := out["segment"]; !ok && out["measure"] == DefaultTimeSeriesMeasure { + out["segment"] = "status" + } + setDefault(out, "chartType", "bar") + } + if err := validateReport(kind, out); err != nil { + return nil, err + } + // A new report must name what it groups or measures. + switch kind { + case KindExecutions: + if err := checkNonEmptyString(out, "groupBy"); err != nil { + return nil, err + } + case KindTimeSeries: + if err := checkNonEmptyString(out, "measure"); err != nil { + return nil, err + } + } + return out, nil +} + +// ValidateReport validates the params of an existing report after an edit, +// returning a new params map. It checks every param that is set and gives +// filters their IDs, but adds no defaults, so an edit changes only what the +// caller named. Unknown keys are preserved (see CheckParamKeys). +func ValidateReport(kind string, params map[string]any) (map[string]any, error) { + if !IsKind(kind) { + return nil, unknownKindError(kind) + } + out := copyParams(params) + if err := validateReport(kind, out); err != nil { + return nil, err + } + return out, nil +} + +// validateReport checks the params that are set, in place: it gives filters +// their IDs and operators, drops empty from/to and an empty segment, and adds +// nothing else. +func validateReport(kind string, out map[string]any) error { + filters, err := ParseFilters(out["filter"]) + if err != nil { + return err + } + for i := range filters { + if filters[i].ID == "" { + filters[i].ID = newID() + } + if filters[i].Operator == "" { + filters[i].Operator = "is" + } + } + out["filter"] = filtersToAny(filters) + + if err := checkEnum(out, "duration", append(slices.Clone(durations), "custom")); err != nil { + return err + } + for _, key := range []string{"from", "to"} { + if err := checkTime(out, key); err != nil { + return err + } + } + + switch kind { + case KindPassFail: + return checkEnum(out, "measure", passFailMeasures) + case KindExecutions: + if err := checkEnum(out, "measure", executionsMeasures); err != nil { + return err + } + return checkString(out, "groupBy") + case KindTimeSeries: + if err := checkString(out, "measure"); err != nil { + return err + } + if err := checkEnum(out, "aggregate", timeSeriesAggregates); err != nil { + return err + } + if err := checkString(out, "segment"); err != nil { + return err + } + if out["segment"] == "" { + delete(out, "segment") + } + if err := checkEnum(out, "chartType", timeSeriesChartTypes); err != nil { + return err + } + if v, ok := out["overlaySuccessRate"]; ok { + if _, isBool := v.(bool); !isBool { + return fmt.Errorf("param overlaySuccessRate must be a boolean, got %T", v) + } + } + } + return nil +} + +func copyParams(params map[string]any) map[string]any { + out := make(map[string]any, len(params)+4) + for k, v := range params { + out[k] = v + } + return out +} + +// MergeParams returns existing params with patch applied on top: keys in patch +// replace the existing ones and a nil value removes a key. +func MergeParams(existing, patch map[string]any) map[string]any { + out := make(map[string]any, len(existing)+len(patch)) + for k, v := range existing { + out[k] = v + } + for k, v := range patch { + if v == nil { + delete(out, k) + continue + } + out[k] = v + } + return out +} + +// ParseFilters decodes params.filter. A missing filter is an empty list. +func ParseFilters(raw any) ([]Filter, error) { + if raw == nil { + return []Filter{}, nil + } + data, err := json.Marshal(raw) + if err != nil { + return nil, fmt.Errorf("param filter is malformed: %w", err) + } + var filters []Filter + if err := json.Unmarshal(data, &filters); err != nil { + return nil, fmt.Errorf("param filter must be an array of {filterConfigurationKey, operator, value} objects: %w", err) + } + for i, f := range filters { + if f.FilterConfigurationKey == "" { + return nil, fmt.Errorf("param filter[%d] is missing filterConfigurationKey", i) + } + if !validFilterValue(f) { + return nil, fmt.Errorf("param filter[%d] (%s) has an invalid value: expected a string or an array of strings; a labels-v2 filter takes an array of selectors with operator 'is', or {labelKey, labelOperator: exists|contains, labelValue} with operator 'where'", i, f.FilterConfigurationKey) + } + } + return filters, nil +} + +// FilterValues gathers every value of the filters with the given key, +// regardless of operator. It is a port of the dashboard's getFilterValue: +// a string value with the "contains" operator becomes "~value", and an +// advanced label filter becomes "key=~value" (contains) or "key" (exists). +func FilterValues(filters []Filter, key string) []string { + result := []string{} + for _, f := range filters { + if f.FilterConfigurationKey != key { + continue + } + if key == FilterLabels { + switch f.Operator { + case "is": + var values []string + if json.Unmarshal(f.Value, &values) == nil { + result = append(result, values...) + } + case "where": + var v LabelFilterValue + if json.Unmarshal(f.Value, &v) != nil { + continue + } + switch v.LabelOperator { + case "contains": + if v.LabelKey != "" && v.LabelValue != "" { + result = append(result, v.LabelKey+"=~"+v.LabelValue) + } + case "exists": + if v.LabelKey != "" { + result = append(result, v.LabelKey) + } + } + } + continue + } + var s string + if json.Unmarshal(f.Value, &s) == nil { + if s == "" { + continue + } + if f.Operator == "contains" { + result = append(result, "~"+s) + } else { + result = append(result, s) + } + continue + } + var values []string + if json.Unmarshal(f.Value, &values) == nil { + result = append(result, values...) + } + } + return result +} + +func validFilterValue(f Filter) bool { + if f.FilterConfigurationKey == FilterLabels { + // Only the two shapes the dashboard writes, which FilterValues reads: + // any other value would pass here and then be silently ignored, and + // the report would query without its label restriction. + switch f.Operator { + case "is", "": + var values []string + return json.Unmarshal(f.Value, &values) == nil + case "where": + var v LabelFilterValue + if json.Unmarshal(f.Value, &v) != nil || v.LabelKey == "" { + return false + } + return v.LabelOperator == "exists" || v.LabelOperator == "contains" + } + return false + } + var s string + if json.Unmarshal(f.Value, &s) == nil { + return true + } + var values []string + return json.Unmarshal(f.Value, &values) == nil +} + +func filtersToAny(filters []Filter) []any { + out := make([]any, 0, len(filters)) + for _, f := range filters { + var value any + _ = json.Unmarshal(f.Value, &value) + out = append(out, map[string]any{ + "filterConfigurationKey": f.FilterConfigurationKey, + "id": f.ID, + "operator": f.Operator, + "value": value, + }) + } + return out +} + +func canonicalFilterKey(key string) string { + if key == "labels" { + return FilterLabels + } + return key +} + +func compact(values []string) []string { + out := make([]string, 0, len(values)) + for _, v := range values { + if v = strings.TrimSpace(v); v != "" { + out = append(out, v) + } + } + return out +} + +func setDefault(params map[string]any, key string, value any) { + if v, ok := params[key]; !ok || v == nil || v == "" { + params[key] = value + } +} + +func checkEnum(params map[string]any, key string, allowed []string) error { + v, ok := params[key] + if !ok || v == "" { + return nil + } + s, isString := v.(string) + if !isString || !slices.Contains(allowed, s) { + return fmt.Errorf("param %s must be one of %s, got %v", key, strings.Join(allowed, ", "), v) + } + return nil +} + +// checkString requires a param, when set, to be a string. +func checkString(params map[string]any, key string) error { + if v, ok := params[key]; ok { + if _, isString := v.(string); !isString { + return fmt.Errorf("param %s must be a string, got %T", key, v) + } + } + return nil +} + +func checkNonEmptyString(params map[string]any, key string) error { + s, ok := params[key].(string) + if !ok || strings.TrimSpace(s) == "" { + return fmt.Errorf("param %s must be a non-empty string", key) + } + return nil +} + +func checkTime(params map[string]any, key string) error { + v, ok := params[key] + if !ok || v == nil || v == "" { + delete(params, key) + return nil + } + s, isString := v.(string) + if !isString { + return fmt.Errorf("param %s must be an RFC3339 timestamp string, got %T", key, v) + } + if _, err := time.Parse(time.RFC3339, s); err != nil { + return fmt.Errorf("param %s must be an RFC3339 timestamp (e.g. 2026-01-15T00:00:00Z): %w", key, err) + } + return nil +} + +func unknownKindError(kind string) error { + return fmt.Errorf("unknown report kind %q (allowed: %s)", kind, strings.Join(Kinds, ", ")) +} + +// newID returns a random identifier for a filter entry; the dashboard uses it +// only as a stable key when rendering the filter list. +func newID() string { + b := make([]byte, 6) + if _, err := rand.Read(b); err != nil { + return fmt.Sprintf("f%d", time.Now().UnixNano()) + } + return hex.EncodeToString(b) +} diff --git a/pkg/mcp/boards/testdata/translation_cases.json b/pkg/mcp/boards/testdata/translation_cases.json new file mode 100644 index 00000000000..9aa5ebbb08f --- /dev/null +++ b/pkg/mcp/boards/testdata/translation_cases.json @@ -0,0 +1,156 @@ +[ + { + "name": "pass-fail default duration is a week ending at the start of tomorrow, and status is dropped", + "kind": "pass-fail", + "now": "2026-09-23T15:04:05Z", + "params": { + "measure": "ratio", + "filter": [ + {"filterConfigurationKey": "workflow", "id": "a", "operator": "is", "value": ["api", ""]}, + {"filterConfigurationKey": "status", "id": "b", "operator": "is", "value": ["failed"]} + ] + }, + "endpoint": "/insights/stats", + "query": {"startDate": "2026-09-17T00:00:00Z", "endDate": "2026-09-24T00:00:00Z", "workflow": "api"} + }, + { + "name": "quarter is 120 days", + "kind": "workflows", + "now": "2026-09-23T15:04:05Z", + "params": {"duration": "quarter"}, + "endpoint": "/insights/workflows", + "query": {"startDate": "2026-05-27T00:00:00Z", "endDate": "2026-09-24T00:00:00Z"} + }, + { + "name": "custom duration without a range falls back to a week", + "kind": "workflows", + "now": "2026-09-23T15:04:05Z", + "params": {"duration": "custom"}, + "endpoint": "/insights/workflows", + "query": {"startDate": "2026-09-17T00:00:00Z", "endDate": "2026-09-24T00:00:00Z"} + }, + { + "name": "from and to are used as given", + "kind": "executions", + "now": "2026-09-23T15:04:05Z", + "params": {"from": "2026-01-01T00:00:00Z", "to": "2026-02-01T00:00:00Z", "duration": "day", "groupBy": "workflow", "measure": "duration"}, + "endpoint": "/insights/executions", + "query": {"startDate": "2026-01-01T00:00:00Z", "endDate": "2026-02-01T00:00:00Z", "groupBy": "workflow"} + }, + { + "name": "only from spans the duration forward", + "kind": "executions", + "now": "2026-09-23T15:04:05Z", + "params": {"from": "2026-01-01T00:00:00Z", "duration": "day"}, + "endpoint": "/insights/executions", + "query": {"startDate": "2026-01-01T00:00:00Z", "endDate": "2026-01-02T00:00:00Z", "groupBy": "status"} + }, + { + "name": "only to spans the duration backward", + "kind": "pass-fail", + "now": "2026-09-23T15:04:05Z", + "params": {"to": "2026-03-31T00:00:00Z", "duration": "month"}, + "endpoint": "/insights/stats", + "query": {"startDate": "2026-03-01T00:00:00Z", "endDate": "2026-03-31T00:00:00Z"} + }, + { + "name": "environment, labels, tags and status filters", + "kind": "workflows", + "now": "2026-09-23T15:04:05Z", + "params": { + "duration": "day", + "filter": [ + {"filterConfigurationKey": "environment", "id": "1", "operator": "is", "value": ["tkcenv_a", "tkcenv_b"]}, + {"filterConfigurationKey": "status", "id": "2", "operator": "is", "value": ["passed", "failed"]}, + {"filterConfigurationKey": "labels-v2", "id": "3", "operator": "is", "value": ["team=core"]}, + {"filterConfigurationKey": "labels-v2", "id": "4", "operator": "where", "value": {"labelKey": "tier", "labelOperator": "contains", "labelValue": "gold"}}, + {"filterConfigurationKey": "labels-v2", "id": "5", "operator": "where", "value": {"labelKey": "owner", "labelOperator": "exists"}}, + {"filterConfigurationKey": "tags", "id": "6", "operator": "is", "value": ["release=1.2,hotfix"]} + ] + }, + "endpoint": "/insights/workflows", + "query": { + "startDate": "2026-09-23T00:00:00Z", + "endDate": "2026-09-24T00:00:00Z", + "env": "tkcenv_a,tkcenv_b", + "status": "passed,failed", + "selector": "team=core,tier=~gold,owner", + "tagFilter": "[\"release=1.2,hotfix\"]" + } + }, + { + "name": "a contains string filter becomes a regex", + "kind": "workflows", + "now": "2026-09-23T15:04:05Z", + "params": { + "duration": "day", + "filter": [{"filterConfigurationKey": "workflow", "id": "1", "operator": "contains", "value": "api-"}] + }, + "endpoint": "/insights/workflows", + "query": {"startDate": "2026-09-23T00:00:00Z", "endDate": "2026-09-24T00:00:00Z", "workflow": "~api-"} + }, + { + "name": "time-series sends measure, aggregate, segment and identity filters", + "kind": "time-series", + "now": "2026-09-23T15:04:05Z", + "params": { + "duration": "day", + "measure": "http_req_duration_p95_ms", + "aggregate": "avg", + "segment": "workflow", + "chartType": "line", + "filter": [ + {"filterConfigurationKey": "workflow", "id": "1", "operator": "is", "value": ["k6"]}, + {"filterConfigurationKey": "scenario", "id": "2", "operator": "is", "value": ["checkout", "~^login"]} + ] + }, + "endpoint": "/insights/series", + "query": { + "startDate": "2026-09-23T00:00:00Z", + "endDate": "2026-09-24T00:00:00Z", + "workflow": "k6", + "measure": "http_req_duration_p95_ms", + "aggregate": "avg", + "segment": "workflow", + "identityFilters": "{\"scenario\":[\"checkout\",\"~^login\"]}" + } + }, + { + "name": "non-base filters are ignored outside time-series", + "kind": "executions", + "now": "2026-09-23T15:04:05Z", + "params": { + "duration": "day", + "filter": [{"filterConfigurationKey": "scenario", "id": "1", "operator": "is", "value": ["checkout"]}] + }, + "endpoint": "/insights/executions", + "query": {"startDate": "2026-09-23T00:00:00Z", "endDate": "2026-09-24T00:00:00Z", "groupBy": "status"} + }, + { + "name": "a zone ahead of UTC ends the range at its own midnight, already the next UTC day", + "kind": "workflows", + "now": "2026-09-23T23:30:00Z", + "timeZone": "Europe/Berlin", + "params": {"duration": "week"}, + "endpoint": "/insights/workflows", + "query": {"startDate": "2026-09-17T22:00:00Z", "endDate": "2026-09-24T22:00:00Z"} + }, + { + "name": "a zone behind UTC ends the range at its own midnight, still the previous UTC day", + "kind": "pass-fail", + "now": "2026-09-24T02:00:00Z", + "timeZone": "America/New_York", + "params": {"duration": "day"}, + "endpoint": "/insights/stats", + "query": {"startDate": "2026-09-23T04:00:00Z", "endDate": "2026-09-24T04:00:00Z"} + }, + { + "name": "a range across a daylight-saving change subtracts fixed minutes, as the dashboard does", + "kind": "workflows", + "now": "2026-11-02T12:00:00Z", + "timeZone": "America/New_York", + "params": {"duration": "week"}, + "endpoint": "/insights/workflows", + "query": {"startDate": "2026-10-27T05:00:00Z", "endDate": "2026-11-03T05:00:00Z"} + } +] diff --git a/pkg/mcp/boards/timezone.go b/pkg/mcp/boards/timezone.go new file mode 100644 index 00000000000..94b0ca049f6 --- /dev/null +++ b/pkg/mcp/boards/timezone.go @@ -0,0 +1,25 @@ +package boards + +import ( + "fmt" + "strings" + "time" + + // Embed the time zone database: the MCP server also runs from minimal + // container images that ship without /usr/share/zoneinfo. + _ "time/tzdata" +) + +// ParseTimeZone resolves an IANA time zone name such as "Europe/Berlin". +// An empty name is UTC. +func ParseTimeZone(name string) (*time.Location, error) { + name = strings.TrimSpace(name) + if name == "" { + return time.UTC, nil + } + loc, err := time.LoadLocation(name) + if err != nil { + return nil, fmt.Errorf("timeZone must be an IANA time zone name such as 'Europe/Berlin' or 'America/New_York': %w", err) + } + return loc, nil +} diff --git a/pkg/mcp/client.go b/pkg/mcp/client.go index 9a3f8f628d6..ae128d9d24a 100644 --- a/pkg/mcp/client.go +++ b/pkg/mcp/client.go @@ -43,4 +43,13 @@ type Client interface { tools.InsightMetricKeysLister tools.InsightMetricSeriesGetter tools.InsightExecutionsLister + + // Insights board interfaces + tools.BoardLister + tools.BoardGetter + tools.BoardSlugChecker + tools.BoardCreator + tools.BoardUpdater + tools.BoardDeleter + tools.BoardInsightQuerier } diff --git a/pkg/mcp/context/debug.go b/pkg/mcp/context/debug.go new file mode 100644 index 00000000000..dd1780c6567 --- /dev/null +++ b/pkg/mcp/context/debug.go @@ -0,0 +1,37 @@ +package mcpcontext + +import ( + "context" +) + +// DebugInfo collects what a tool call did, for the debug output of the +// call. It is not safe for concurrent use: a tool that makes several client +// calls in parallel gives each its own DebugInfo (see WithDebugInfo) and +// merges them once they are done. +type DebugInfo struct { + Source string `json:"source"` // "http", "file", "database", "cache", etc. + Data map[string]any `json:"data"` // Source-specific debug data +} + +func NewDebugInfo() *DebugInfo { + return &DebugInfo{ + Data: make(map[string]any), + } +} + +const debugInfoKey contextKey = "debug_info" + +// WithDebugInfo returns a context carrying a new DebugInfo, and that DebugInfo. +func WithDebugInfo(ctx context.Context) (context.Context, *DebugInfo) { + debugInfo := NewDebugInfo() + newCtx := context.WithValue(ctx, debugInfoKey, debugInfo) + return newCtx, debugInfo +} + +// GetDebugInfo returns the DebugInfo in ctx, or nil when debugging is off. +func GetDebugInfo(ctx context.Context) *DebugInfo { + if debugInfo, ok := ctx.Value(debugInfoKey).(*DebugInfo); ok { + return debugInfo + } + return nil +} diff --git a/pkg/mcp/debug.go b/pkg/mcp/debug.go index 0fdd476be41..8c77f873ea6 100644 --- a/pkg/mcp/debug.go +++ b/pkg/mcp/debug.go @@ -2,32 +2,22 @@ package mcp import ( "context" + + mcpcontext "github.com/kubeshop/testkube/pkg/mcp/context" ) -type DebugInfo struct { - Source string `json:"source"` // "http", "file", "database", "cache", etc. - Data map[string]any `json:"data"` // Source-specific debug data -} +// DebugInfo lives in mcpcontext so the tools, which this package imports, +// can use it too. These keep the existing names working. +type DebugInfo = mcpcontext.DebugInfo func NewDebugInfo() *DebugInfo { - return &DebugInfo{ - Data: make(map[string]any), - } + return mcpcontext.NewDebugInfo() } -type contextKey string - -const debugInfoKey contextKey = "debug_info" - func WithDebugInfo(ctx context.Context) (context.Context, *DebugInfo) { - debugInfo := NewDebugInfo() - newCtx := context.WithValue(ctx, debugInfoKey, debugInfo) - return newCtx, debugInfo + return mcpcontext.WithDebugInfo(ctx) } func GetDebugInfo(ctx context.Context) *DebugInfo { - if debugInfo, ok := ctx.Value(debugInfoKey).(*DebugInfo); ok { - return debugInfo - } - return nil + return mcpcontext.GetDebugInfo(ctx) } diff --git a/pkg/mcp/formatters/boards.go b/pkg/mcp/formatters/boards.go new file mode 100644 index 00000000000..b7093f62aa7 --- /dev/null +++ b/pkg/mcp/formatters/boards.go @@ -0,0 +1,289 @@ +package formatters + +import ( + "encoding/json" + "fmt" + "sort" + + "github.com/kubeshop/testkube/pkg/mcp/boards" +) + +// maxRenderedWorkflows bounds the workflows report, which lists every workflow +// in range and can be large. +const maxRenderedWorkflows = 25 + +// --- list_boards ------------------------------------------------------------- + +type boardSummaryInput struct { + ID string `json:"id"` + Slug string `json:"slug"` + Name string `json:"name"` + Description string `json:"description"` + UpdatedAt string `json:"updatedAt"` + Creator string `json:"creator"` + Shared bool `json:"shared"` + IsUserFavorite bool `json:"isUserFavorite"` + IsOrgFavorite bool `json:"isOrgFavorite"` + AnalysisCount int `json:"analysisCount"` +} + +type formattedBoardSummary struct { + ID string `json:"id"` + Slug string `json:"slug"` + Name string `json:"name"` + Description string `json:"description,omitempty"` + Shared bool `json:"shared"` + Reports int `json:"reports"` + Pinned bool `json:"pinned,omitempty"` + OrgPinned bool `json:"orgPinned,omitempty"` + UpdatedAt string `json:"updatedAt,omitempty"` +} + +// FormatBoardList compacts a board list. The Control Plane reports the total +// in a response header, which the tool does not see, so hasMore is inferred +// from a full page. +func FormatBoardList(raw string, page, pageSize int) (string, error) { + input, isEmpty, err := ParseJSON[[]boardSummaryInput](raw) + if err != nil { + return "", err + } + if isEmpty || len(input) == 0 { + if page > 0 { + return "No more boards.", nil + } + return "No boards found.", nil + } + + result := make([]formattedBoardSummary, 0, len(input)) + for _, b := range input { + result = append(result, formattedBoardSummary{ + ID: b.ID, + Slug: b.Slug, + Name: b.Name, + Description: b.Description, + Shared: b.Shared, + Reports: b.AnalysisCount, + Pinned: b.IsUserFavorite, + OrgPinned: b.IsOrgFavorite, + UpdatedAt: b.UpdatedAt, + }) + } + + return FormatJSON(struct { + Boards []formattedBoardSummary `json:"boards"` + Page int `json:"page"` + HasMore bool `json:"hasMore"` + }{Boards: result, Page: page, HasMore: pageSize > 0 && len(input) >= pageSize}) +} + +// --- get_board --------------------------------------------------------------- + +// FormattedBoard is the compact view of a board the board tools return. +type FormattedBoard struct { + ID string `json:"id"` + Slug string `json:"slug"` + Name string `json:"name"` + Description string `json:"description,omitempty"` + Shared bool `json:"shared"` + Creator string `json:"creator,omitempty"` + UpdatedAt string `json:"updatedAt,omitempty"` + Reports []boards.Report `json:"reports"` + Layout [][]string `json:"layout"` + Unplaced []string `json:"unplaced,omitempty"` +} + +// BoardView builds the compact view of a board: reports in the order the +// dashboard shows them, the layout as rows of report IDs, and the reports the +// layout leaves out (which the dashboard does not show). +func BoardView(b *boards.Board) (FormattedBoard, error) { + ordered, unplaced, err := b.OrderedReports() + if err != nil { + return FormattedBoard{}, err + } + layout, err := boards.ParseLayout(b.Layout) + if err != nil { + return FormattedBoard{}, err + } + if ordered == nil { + ordered = []boards.Report{} + } + return FormattedBoard{ + ID: b.ID, + Slug: b.Slug, + Name: b.Name, + Description: b.Description, + Shared: b.Shared, + Creator: b.Creator, + UpdatedAt: b.UpdatedAt, + Reports: ordered, + Layout: layout.RowIDs(), + Unplaced: unplaced, + }, nil +} + +// FormatBoard formats a board response. +func FormatBoard(raw string) (string, error) { + b, err := boards.ParseBoard(raw) + if err != nil { + return "", err + } + view, err := BoardView(b) + if err != nil { + return "", err + } + return FormatJSON(view) +} + +// --- render_board ------------------------------------------------------------ + +type statsSeriesInput struct { + Total float64 `json:"total"` + Values []json.RawMessage `json:"values"` +} + +type passFailInput struct { + RatioStats statsSeriesInput `json:"ratioStats"` + TotalStats statsSeriesInput `json:"totalStats"` + FailedStats statsSeriesInput `json:"failedStats"` +} + +type groupedInput struct { + Total float64 `json:"total"` + Values []json.RawMessage `json:"values"` +} + +type executionsInput struct { + Count groupedInput `json:"count"` + Duration groupedInput `json:"duration"` +} + +type workflowSummaryInput struct { + Name string `json:"name"` + TotalExecutionCount int `json:"totalExecutionCount"` + FailedExecutionCount int `json:"failedExecutionCount"` + AverageDuration int64 `json:"averageDuration"` + P95Duration int64 `json:"p95Duration"` + LastRunAt string `json:"lastRunAt"` +} + +type formattedWorkflowSummary struct { + Name string `json:"name"` + Executions int `json:"executions"` + Failed int `json:"failed"` + PassRate float64 `json:"passRate"` + AvgDurationMs int64 `json:"avgDurationMs"` + P95DurationMs int64 `json:"p95DurationMs"` + LastRunAt string `json:"lastRunAt,omitempty"` +} + +// ReportData summarizes the response of the query that renders a report, for +// the value that report shows. measure is the report's params.measure; an +// empty measure means the kind's default. +func ReportData(kind, measure, raw string, maxSamples int) (any, error) { + if maxSamples <= 0 { + maxSamples = defaultInsightSeriesSamples + } + switch kind { + case boards.KindPassFail: + input, _, err := ParseJSON[passFailInput](raw) + if err != nil { + return nil, err + } + selected := input.RatioStats + switch measure { + case "failed-count": + selected = input.FailedStats + case "total-count": + selected = input.TotalStats + default: + measure = "ratio" + } + points := tuples(selected.Values) + return map[string]any{ + "measure": measure, + "total": selected.Total, + "passRatio": input.RatioStats.Total, + "totalExecutions": input.TotalStats.Total, + "failedExecutions": input.FailedStats.Total, + "points": len(points), + "samples": downsampleInsightPoints(points, maxSamples), + }, nil + + case boards.KindExecutions: + input, _, err := ParseJSON[executionsInput](raw) + if err != nil { + return nil, err + } + selected := input.Count + if measure == "duration" { + selected = input.Duration + } else { + measure = "count" + } + groups := make([]map[string]any, 0, len(selected.Values)) + for _, t := range tuples(selected.Values) { + groups = append(groups, map[string]any{"group": t[0], "value": t[1]}) + } + unit := "executions" + if measure == "duration" { + unit = "avg ms" + } + return map[string]any{"measure": measure, "unit": unit, "total": selected.Total, "groups": groups}, nil + + case boards.KindWorkflows: + input, _, err := ParseJSON[[]workflowSummaryInput](raw) + if err != nil { + return nil, err + } + sort.SliceStable(input, func(i, j int) bool { + return input[i].TotalExecutionCount > input[j].TotalExecutionCount + }) + count := len(input) + if len(input) > maxRenderedWorkflows { + input = input[:maxRenderedWorkflows] + } + workflows := make([]formattedWorkflowSummary, 0, len(input)) + for _, w := range input { + passRate := 0.0 + if w.TotalExecutionCount > 0 { + passRate = float64(w.TotalExecutionCount-w.FailedExecutionCount) / float64(w.TotalExecutionCount) * 100 + } + workflows = append(workflows, formattedWorkflowSummary{ + Name: w.Name, + Executions: w.TotalExecutionCount, + Failed: w.FailedExecutionCount, + PassRate: passRate, + AvgDurationMs: w.AverageDuration, + P95DurationMs: w.P95Duration, + LastRunAt: w.LastRunAt, + }) + } + return map[string]any{"workflowCount": count, "truncated": count > len(workflows), "workflows": workflows}, nil + + case boards.KindTimeSeries: + data, _, err := ParseJSON[[]insightSeriesDatum](raw) + if err != nil { + return nil, err + } + total, segments := summarizeInsightSeries(data, maxSamples) + if segments == nil { + segments = []formattedInsightSegment{} + } + return map[string]any{"pointCount": total, "series": segments}, nil + } + return nil, fmt.Errorf("unknown report kind %q", kind) +} + +// tuples decodes the [label, value] pairs the reporting endpoints return. +// Pairs that are not two elements long are skipped. +func tuples(values []json.RawMessage) [][2]any { + out := make([][2]any, 0, len(values)) + for _, v := range values { + var pair []any + if json.Unmarshal(v, &pair) != nil || len(pair) != 2 { + continue + } + out = append(out, [2]any{pair[0], pair[1]}) + } + return out +} diff --git a/pkg/mcp/formatters/boards_test.go b/pkg/mcp/formatters/boards_test.go new file mode 100644 index 00000000000..d3f1b5f564b --- /dev/null +++ b/pkg/mcp/formatters/boards_test.go @@ -0,0 +1,85 @@ +package formatters + +import ( + "encoding/json" + "fmt" + "strings" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/kubeshop/testkube/pkg/mcp/boards" +) + +func TestFormatBoardList(t *testing.T) { + out, err := FormatBoardList(`[{"id":"b1","slug":"s","name":"N","shared":false,"analysisCount":2,"isUserFavorite":true,"layoutSummary":{"rows":1,"cells":2}}]`, 0, 1) + require.NoError(t, err) + assert.JSONEq(t, `{"boards":[{"id":"b1","slug":"s","name":"N","shared":false,"reports":2,"pinned":true}],"page":0,"hasMore":true}`, out) + + out, err = FormatBoardList(`[]`, 0, 20) + require.NoError(t, err) + assert.Equal(t, "No boards found.", out) +} + +func TestReportData(t *testing.T) { + t.Run("pass-fail selects the measure", func(t *testing.T) { + raw := `{"ratioStats":{"total":75.5,"values":[["2026-09-01",50],["2026-09-02",100]]}, + "totalStats":{"total":8,"values":[["2026-09-01",4],["2026-09-02",4]]}, + "failedStats":{"total":2,"values":[["2026-09-01",2],["2026-09-02",0]]}}` + data, err := ReportData(boards.KindPassFail, "failed-count", raw, 0) + require.NoError(t, err) + m := data.(map[string]any) + assert.Equal(t, "failed-count", m["measure"]) + assert.Equal(t, float64(2), m["total"]) + assert.Equal(t, float64(75.5), m["passRatio"]) + assert.Equal(t, 2, m["points"]) + }) + + t.Run("pass-fail downsamples", func(t *testing.T) { + values := make([]string, 0, 100) + for i := 0; i < 100; i++ { + values = append(values, fmt.Sprintf(`["d%d",%d]`, i, i)) + } + raw := `{"ratioStats":{"total":1,"values":[` + strings.Join(values, ",") + `]},"totalStats":{"total":0,"values":[]},"failedStats":{"total":0,"values":[]}}` + data, err := ReportData(boards.KindPassFail, "", raw, 10) + require.NoError(t, err) + m := data.(map[string]any) + assert.Equal(t, "ratio", m["measure"]) + assert.Equal(t, 100, m["points"]) + assert.Len(t, m["samples"], 10) + }) + + t.Run("executions groups", func(t *testing.T) { + raw := `{"count":{"total":5,"values":[["passed",4],["failed",1]]},"duration":{"total":900,"values":[["passed",100],["failed",800]]}}` + data, err := ReportData(boards.KindExecutions, "duration", raw, 0) + require.NoError(t, err) + out, _ := json.Marshal(data) + assert.JSONEq(t, `{"measure":"duration","unit":"avg ms","total":900,"groups":[{"group":"passed","value":100},{"group":"failed","value":800}]}`, string(out)) + }) + + t.Run("workflows are ranked and truncated", func(t *testing.T) { + items := make([]string, 0, 30) + for i := 0; i < 30; i++ { + items = append(items, fmt.Sprintf(`{"name":"wf%d","totalExecutionCount":%d,"failedExecutionCount":1,"averageDuration":10,"p95Duration":20,"lastRunAt":"2026-09-01T00:00:00Z"}`, i, i+1)) + } + data, err := ReportData(boards.KindWorkflows, "", "["+strings.Join(items, ",")+"]", 0) + require.NoError(t, err) + m := data.(map[string]any) + assert.Equal(t, 30, m["workflowCount"]) + assert.Equal(t, true, m["truncated"]) + workflows := m["workflows"].([]formattedWorkflowSummary) + require.Len(t, workflows, maxRenderedWorkflows) + assert.Equal(t, "wf29", workflows[0].Name) + assert.InDelta(t, 96.67, workflows[0].PassRate, 0.01) + }) + + t.Run("time-series summarizes segments", func(t *testing.T) { + raw := `[{"ts":1,"value":3,"segments":[{"label":"passed","value":2},{"label":"failed","value":1}]},{"ts":2,"value":0}]` + data, err := ReportData(boards.KindTimeSeries, "execution-count", raw, 0) + require.NoError(t, err) + m := data.(map[string]any) + assert.Equal(t, 2, m["pointCount"]) + assert.Len(t, m["series"], 2) + }) +} diff --git a/pkg/mcp/formatters/insights.go b/pkg/mcp/formatters/insights.go index 060ae32b499..be5e542d0f4 100644 --- a/pkg/mcp/formatters/insights.go +++ b/pkg/mcp/formatters/insights.go @@ -120,6 +120,17 @@ func FormatInsightMetricSeries(raw string, maxSamples int) (string, error) { if isEmpty || len(data) == 0 { return "No metric data found for the given query.", nil } + + total, segments := summarizeInsightSeries(data, maxSamples) + return FormatJSON(struct { + PointCount int `json:"pointCount"` + Series []formattedInsightSegment `json:"series"` + }{PointCount: total, Series: segments}) +} + +// summarizeInsightSeries groups a time series by segment and summarizes each +// with downsampled points. maxSamples <= 0 uses the default. +func summarizeInsightSeries(data []insightSeriesDatum, maxSamples int) (int, []formattedInsightSegment) { if maxSamples <= 0 { maxSamples = defaultInsightSeriesSamples } @@ -189,15 +200,12 @@ func FormatInsightMetricSeries(raw string, maxSamples int) (string, error) { }) } - return FormatJSON(struct { - PointCount int `json:"pointCount"` - Series []formattedInsightSegment `json:"series"` - }{PointCount: total, Series: segments}) + return total, segments } // downsampleInsightPoints evenly reduces a series to at most maxPoints. // For maxPoints >= 2 it always keeps the first and last point. -func downsampleInsightPoints(values []insightSeriesPoint, maxPoints int) []insightSeriesPoint { +func downsampleInsightPoints[T any](values []T, maxPoints int) []T { if maxPoints <= 0 || len(values) <= maxPoints { return values } @@ -205,7 +213,7 @@ func downsampleInsightPoints(values []insightSeriesPoint, maxPoints int) []insig return values[len(values)-1:] } - result := make([]insightSeriesPoint, 0, maxPoints) + result := make([]T, 0, maxPoints) result = append(result, values[0]) step := float64(len(values)-1) / float64(maxPoints-1) for i := 1; i < maxPoints-1; i++ { diff --git a/pkg/mcp/server.go b/pkg/mcp/server.go index 638f8173671..7c9a947a57f 100644 --- a/pkg/mcp/server.go +++ b/pkg/mcp/server.go @@ -105,6 +105,18 @@ func NewMCPServer(cfg MCPServerConfig, client Client) (*server.MCPServer, error) mcpServer.AddTool(tools.GetInsightMetricSeries(client)) mcpServer.AddTool(tools.ListInsightExecutions(client)) + // Insights board tools. Registered unconditionally, like the insight tools: + // the board routes do not answer HEAD, so SupportsEndpoint cannot probe them. + mcpServer.AddTool(tools.ListBoards(client)) + mcpServer.AddTool(tools.GetBoard(client)) + mcpServer.AddTool(tools.CreateBoard(client)) + mcpServer.AddTool(tools.UpdateBoard(client)) + mcpServer.AddTool(tools.AddBoardReport(client)) + mcpServer.AddTool(tools.UpdateBoardReport(client)) + mcpServer.AddTool(tools.RemoveBoardReport(client)) + mcpServer.AddTool(tools.DeleteBoard(client)) + mcpServer.AddTool(tools.RenderBoard(client)) + return mcpServer, nil } diff --git a/pkg/mcp/tools/boards.go b/pkg/mcp/tools/boards.go new file mode 100644 index 00000000000..8095972d77b --- /dev/null +++ b/pkg/mcp/tools/boards.go @@ -0,0 +1,972 @@ +package tools + +import ( + "context" + "encoding/json" + "errors" + "fmt" + "slices" + "strings" + "sync" + "time" + + "github.com/mark3labs/mcp-go/mcp" + "github.com/mark3labs/mcp-go/server" + + "github.com/kubeshop/testkube/pkg/mcp/boards" + mcpcontext "github.com/kubeshop/testkube/pkg/mcp/context" + "github.com/kubeshop/testkube/pkg/mcp/formatters" +) + +// Board tools manage Insights boards: saved, organization-wide dashboards made +// of reports (charts) over execution data. Boards are organization-scoped, not +// environment-scoped; a report narrows itself to environments through its own +// filter. The Control Plane only serves boards to signed-in users, so these +// tools fail for API tokens (see ErrBoardsRequireUser). +// +// Report params are an opaque object to the Control Plane. The tools validate +// and normalize them with pkg/mcp/boards so the dashboard can render and edit +// what they write. + +// ErrBoardsRequireUser is returned by a client when the session authenticates +// with an API token, which the Control Plane refuses for every board endpoint. +var ErrBoardsRequireUser = errors.New("insights boards require a signed-in user session; API tokens are not supported. " + + "Sign in with `testkube login` and restart the MCP server, or connect to the hosted MCP endpoint with your user account. " + + "In Docker or environment-variable mode, TK_ACCESS_TOKEN must be a user access token, not an API token (tkcapi_...)") + +// ErrBoardChanged is returned by a client when a conditional board update is +// refused because the board no longer has the version the update carries in +// ExpectedVersion. +var ErrBoardChanged = errors.New("the board changed since it was read") + +const renderConcurrency = 4 + +// boardWriteAttempts bounds how many times a board write is rebuilt from a +// fresh read after losing a race with a concurrent edit. +const boardWriteAttempts = 3 + +// ListBoardsParams filters the board list. +type ListBoardsParams struct { + Name string + Private bool + Shared bool + UserFavorite bool + OrgFavorite bool + Page int + PageSize int +} + +// CreateBoardParams describes a new board. +type CreateBoardParams struct { + Name string `json:"name"` + Slug string `json:"slug,omitempty"` + Description string `json:"description,omitempty"` + IsPrivate bool `json:"isPrivate"` +} + +// BoardContentPatch adds, replaces or removes one report of a board. +type BoardContentPatch struct { + Action string `json:"action"` // create | update | delete + ContentKind string `json:"content_kind"` + ContentID string `json:"content_id,omitempty"` + ContentData *boards.ReportDraft `json:"content_data,omitempty"` +} + +// UpdateBoardRequest is the body of a board update. Nil fields are left +// unchanged. The board tools always send Description too: current Control +// Planes keep a description an update omits, but older ones clear it. +// +// Resending a value read earlier would overwrite a concurrent edit of it, so +// every board write also sends ExpectedVersion, the version of the board it +// read. The Control Plane then refuses the write with 409 if the board has +// changed since, and the tools read it again and rebuild the write. A Control +// Plane that predates the field ignores it. +type UpdateBoardRequest struct { + Name *string `json:"name,omitempty"` + Description *string `json:"description,omitempty"` + Slug *string `json:"slug,omitempty"` + IsPrivate *bool `json:"isPrivate,omitempty"` + Layout json.RawMessage `json:"layout,omitempty"` + Content *BoardContentPatch `json:"content,omitempty"` + ExpectedVersion *int64 `json:"expectedVersion,omitempty"` +} + +// BoardLister lists the boards visible to the user. +type BoardLister interface { + ListBoards(ctx context.Context, params ListBoardsParams) (string, error) +} + +// BoardGetter returns a board by ID or slug. +type BoardGetter interface { + GetBoard(ctx context.Context, board string) (string, error) +} + +// BoardSlugChecker reports whether a board slug is still free. +type BoardSlugChecker interface { + CheckBoardSlug(ctx context.Context, slug string) (bool, error) +} + +// BoardCreator creates a board. +type BoardCreator interface { + CreateBoard(ctx context.Context, params CreateBoardParams) (string, error) +} + +// BoardUpdater updates a board's details, layout or reports. +type BoardUpdater interface { + UpdateBoard(ctx context.Context, board string, request UpdateBoardRequest) (string, error) +} + +// BoardDeleter deletes a board. +type BoardDeleter interface { + DeleteBoard(ctx context.Context, board string) error +} + +// BoardInsightQuerier runs the insight query that renders a board report and +// returns the raw response. +type BoardInsightQuerier interface { + QueryBoardInsights(ctx context.Context, query boards.InsightQuery) (string, error) +} + +// BoardEditor reads a board and then writes it. +type BoardEditor interface { + BoardGetter + BoardUpdater +} + +// BoardCreatorWithSlugCheck creates a board after checking its slug. +type BoardCreatorWithSlugCheck interface { + BoardCreator + BoardSlugChecker +} + +// BoardRemover reads a board and then deletes it. +type BoardRemover interface { + BoardGetter + BoardDeleter +} + +// BoardRenderer reads a board and runs its reports' queries. +type BoardRenderer interface { + BoardGetter + BoardInsightQuerier +} + +// ListBoards creates a tool for listing Insights boards. +func ListBoards(client BoardLister) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("list_boards", + mcp.WithDescription(ListBoardsDescription), + mcp.WithReadOnlyHintAnnotation(true), + mcp.WithString("name", mcp.Description("Filter boards whose name contains this text.")), + mcp.WithString("visibility", mcp.Description("'all' (default), 'private' (only your private boards) or 'shared' (only boards shared with the organization)."), mcp.Enum("all", "private", "shared")), + mcp.WithString("favorite", mcp.Description("'user' for boards you pinned, 'org' for boards pinned for the organization."), mcp.Enum("user", "org")), + mcp.WithNumber("page", mcp.Description(PageDescription)), + mcp.WithNumber("pageSize", mcp.Description("Number of boards per page (default: 20, max: 100)")), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + name, err := OptionalParam[string](request, "name") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + visibility, err := OptionalParam[string](request, "visibility") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + favorite, err := OptionalParam[string](request, "favorite") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + page, err := OptionalIntParam(request, "page") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + pageSize, err := OptionalIntParamWithDefault(request, "pageSize", 20) + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + if page < 0 { + page = 0 + } + if pageSize < 1 || pageSize > 100 { + pageSize = min(max(pageSize, 1), 100) + } + + params := ListBoardsParams{Name: name, Page: page, PageSize: pageSize} + switch visibility { + case "", "all": + case "private": + params.Private = true + case "shared": + params.Shared = true + default: + return mcp.NewToolResultError("visibility must be one of all, private, shared"), nil + } + switch favorite { + case "": + case "user": + params.UserFavorite = true + case "org": + params.OrgFavorite = true + default: + return mcp.NewToolResultError("favorite must be one of user, org"), nil + } + + result, err := client.ListBoards(ctx, params) + if err != nil { + return boardError("list boards", err), nil + } + formatted, err := formatters.FormatBoardList(result, page, pageSize) + if err != nil { + return mcp.NewToolResultError(fmt.Sprintf("Failed to format boards: %v", err)), nil + } + return mcp.NewToolResultText(formatted), nil + } + + return tool, handler +} + +// GetBoard creates a tool for viewing a board and its reports. +func GetBoard(client BoardGetter) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("get_board", + mcp.WithDescription(GetBoardDescription), + mcp.WithReadOnlyHintAnnotation(true), + mcp.WithString("board", mcp.Required(), mcp.Description(BoardIdDescription)), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + board, err := RequiredParam[string](request, "board") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + result, err := client.GetBoard(ctx, board) + if err != nil { + return boardError("get board", err), nil + } + formatted, err := formatters.FormatBoard(result) + if err != nil { + return mcp.NewToolResultError(fmt.Sprintf("Failed to format board: %v", err)), nil + } + return mcp.NewToolResultText(formatted), nil + } + + return tool, handler +} + +// CreateBoard creates a tool for creating an empty board. +func CreateBoard(client BoardCreatorWithSlugCheck) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("create_board", + mcp.WithDescription(CreateBoardDescription), + mcp.WithString("name", mcp.Required(), mcp.Description("Name of the board.")), + mcp.WithString("description", mcp.Description("Description of the board.")), + mcp.WithString("slug", mcp.Description("URL-friendly identifier. Generated from the name when omitted; must be unused in the organization.")), + mcp.WithBoolean("private", mcp.Description("Create the board private to you instead of shared with the organization (default: false, shared).")), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + name, err := RequiredParam[string](request, "name") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + description, err := OptionalParam[string](request, "description") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + slug, err := OptionalParam[string](request, "slug") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + private, err := OptionalParam[bool](request, "private") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + + // The Control Plane does not check a slug it is given on create. + if slug = strings.TrimSpace(slug); slug != "" { + available, err := client.CheckBoardSlug(ctx, slug) + if err != nil { + return boardError("check board slug", err), nil + } + if !available { + return mcp.NewToolResultError(fmt.Sprintf("slug %q is already used by another board; choose another or omit it to generate one", slug)), nil + } + } + + result, err := client.CreateBoard(ctx, CreateBoardParams{Name: name, Slug: slug, Description: description, IsPrivate: private}) + if err != nil { + return boardError("create board", err), nil + } + formatted, err := formatters.FormatBoard(result) + if err != nil { + return mcp.NewToolResultError(fmt.Sprintf("Failed to format board: %v", err)), nil + } + return mcp.NewToolResultText(formatted), nil + } + + return tool, handler +} + +// UpdateBoard creates a tool for changing a board's details or layout. +func UpdateBoard(client BoardEditor) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("update_board", + mcp.WithDescription(UpdateBoardDescription), + mcp.WithString("board", mcp.Required(), mcp.Description(BoardIdDescription)), + mcp.WithString("name", mcp.Description("New name.")), + mcp.WithString("description", mcp.Description("New description. Pass an empty string to clear it; omit to keep it.")), + mcp.WithString("slug", mcp.Description("New URL-friendly identifier; must be unused in the organization.")), + mcp.WithBoolean("private", mcp.Description("true makes the board private to its creator, false shares it with the organization. Only the creator or an organization admin can change this.")), + mcp.WithObject("layout", mcp.Description(BoardLayoutDescription)), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + boardRef, err := RequiredParam[string](request, "board") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + // The arguments: every field the caller did not name keeps its value. + var change UpdateBoardRequest + var layout *boards.Layout + if name, ok, err := OptionalParamOK[string](request, "name"); err != nil { + return mcp.NewToolResultError(err.Error()), nil + } else if ok && strings.TrimSpace(name) != "" { + change.Name = &name + } + if description, ok, err := OptionalParamOK[string](request, "description"); err != nil { + return mcp.NewToolResultError(err.Error()), nil + } else if ok { + change.Description = &description + } + if slug, ok, err := OptionalParamOK[string](request, "slug"); err != nil { + return mcp.NewToolResultError(err.Error()), nil + } else if ok && strings.TrimSpace(slug) != "" { + slug = strings.TrimSpace(slug) + change.Slug = &slug + } + if private, ok, err := OptionalParamOK[bool](request, "private"); err != nil { + return mcp.NewToolResultError(err.Error()), nil + } else if ok { + change.IsPrivate = &private + } + if raw, ok := request.GetArguments()["layout"]; ok && raw != nil { + parsed, err := parseLayoutParam(raw) + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + layout = &parsed + } + if change.Name == nil && change.Description == nil && change.Slug == nil && change.IsPrivate == nil && layout == nil { + return mcp.NewToolResultError("nothing to update: pass at least one of name, description, slug, private, layout"), nil + } + + result, _, errResult := writeBoard(ctx, client, boardRef, "update board", func(board *boards.Board) (UpdateBoardRequest, *mcp.CallToolResult) { + req := change + if req.Description == nil { + req.Description = &board.Description + } + if layout != nil { + // Checked against the board as it is now, so a report added + // since the caller read the board is not silently unplaced. + if err := layout.Validate(board.ReportIDs()); err != nil { + return req, mcp.NewToolResultError(err.Error()) + } + req.Layout, _ = json.Marshal(layout) + } + return req, nil + }) + if errResult != nil { + return errResult, nil + } + formatted, err := formatters.FormatBoard(result) + if err != nil { + return mcp.NewToolResultError(fmt.Sprintf("Failed to format board: %v", err)), nil + } + return mcp.NewToolResultText(formatted), nil + } + + return tool, handler +} + +// AddBoardReport creates a tool for adding a report to a board. +func AddBoardReport(client BoardEditor) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("add_board_report", + mcp.WithDescription(AddBoardReportDescription), + mcp.WithString("board", mcp.Required(), mcp.Description(BoardIdDescription)), + mcp.WithString("kind", mcp.Required(), mcp.Description(BoardReportKindDescription), mcp.Enum(boards.Kinds...)), + mcp.WithString("name", mcp.Required(), mcp.Description("Title of the report.")), + mcp.WithString("description", mcp.Description("Description of the report.")), + mcp.WithObject("params", mcp.Description(BoardReportParamsDescription)), + mcp.WithObject("filters", mcp.Description(BoardReportFiltersDescription)), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + boardRef, err := RequiredParam[string](request, "board") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + kind, err := RequiredParam[string](request, "kind") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + name, err := RequiredParam[string](request, "name") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + description, err := OptionalParam[string](request, "description") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + params, filters, err := reportInput(request, kind) + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + if params == nil { + params = map[string]any{} + } + if err := boards.ApplyFilters(params, filters); err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + normalized, err := boards.NormalizeReport(kind, params) + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + + result, board, errResult := writeBoard(ctx, client, boardRef, "add report", func(board *boards.Board) (UpdateBoardRequest, *mcp.CallToolResult) { + return UpdateBoardRequest{ + Description: &board.Description, + Content: &BoardContentPatch{ + Action: "create", + ContentKind: "report", + ContentData: &boards.ReportDraft{Kind: kind, Name: name, Description: description, Params: normalized}, + }, + }, nil + }) + if errResult != nil { + return errResult, nil + } + before := board.ReportIDs() + return boardChangeResult(result, func(updated *boards.Board) string { + for _, id := range updated.ReportIDs() { + if !slices.Contains(before, id) { + return id + } + } + return "" + }) + } + + return tool, handler +} + +// UpdateBoardReport creates a tool for changing a report on a board. +func UpdateBoardReport(client BoardEditor) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("update_board_report", + mcp.WithDescription(UpdateBoardReportDescription), + mcp.WithString("board", mcp.Required(), mcp.Description(BoardIdDescription)), + mcp.WithString("reportId", mcp.Required(), mcp.Description("ID of the report to change (from get_board).")), + mcp.WithString("name", mcp.Description("New title.")), + mcp.WithString("description", mcp.Description("New description. Pass an empty string to clear it; omit to keep it.")), + mcp.WithString("kind", mcp.Description("New report kind. Changing it starts from that kind's defaults instead of the old params."), mcp.Enum(boards.Kinds...)), + mcp.WithObject("params", mcp.Description(BoardReportParamsDescription+" Merged into the current params unless replaceParams is true; a null value removes a param.")), + mcp.WithObject("filters", mcp.Description(BoardReportFiltersDescription+" Each key given replaces that key's current filters; an empty list removes them.")), + mcp.WithBoolean("replaceParams", mcp.Description("Replace the params entirely instead of merging (default: false).")), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + boardRef, err := RequiredParam[string](request, "board") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + reportID, err := RequiredParam[string](request, "reportId") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + replace, err := OptionalParam[bool](request, "replaceParams") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + + newKind, err := OptionalParam[string](request, "kind") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + newName, err := OptionalParam[string](request, "name") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + newDescription, hasDescription, err := OptionalParamOK[string](request, "description") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + + result, _, errResult := writeBoard(ctx, client, boardRef, "update report", func(board *boards.Board) (UpdateBoardRequest, *mcp.CallToolResult) { + // The Control Plane silently ignores an update of an unknown report. + existing, ok := board.FindReport(reportID) + if !ok { + return UpdateBoardRequest{}, reportNotFound(board, reportID) + } + + // The merge starts from the report as it is now, so a concurrent + // edit of params the caller did not name is kept. + draft := boards.ReportDraft{Kind: existing.Kind, Name: existing.Name, Description: existing.Description} + replaceParams := replace + if newKind != "" && newKind != existing.Kind { + draft.Kind = newKind + replaceParams = true + } + if strings.TrimSpace(newName) != "" { + draft.Name = newName + } + if hasDescription { + draft.Description = newDescription + } + + patch, filters, err := reportInput(request, draft.Kind) + if err != nil { + return UpdateBoardRequest{}, mcp.NewToolResultError(err.Error()) + } + base := existing.Params + if replaceParams { + base = nil + } + params := boards.MergeParams(base, patch) + if err := boards.ApplyFilters(params, filters); err != nil { + return UpdateBoardRequest{}, mcp.NewToolResultError(err.Error()) + } + // A new kind or replaced params start from the kind's defaults, as a + // new report does. An edit only validates: filling defaults into the + // stored params could change what the report shows, e.g. its window. + normalize := boards.ValidateReport + if replaceParams { + normalize = boards.NormalizeReport + } + if draft.Params, err = normalize(draft.Kind, params); err != nil { + return UpdateBoardRequest{}, mcp.NewToolResultError(err.Error()) + } + + return UpdateBoardRequest{ + Description: &board.Description, + Content: &BoardContentPatch{ + Action: "update", + ContentKind: "report", + ContentID: reportID, + ContentData: &draft, + }, + }, nil + }) + if errResult != nil { + return errResult, nil + } + return boardChangeResult(result, func(*boards.Board) string { return reportID }) + } + + return tool, handler +} + +// RemoveBoardReport creates a tool for removing a report from a board. +func RemoveBoardReport(client BoardEditor) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("remove_board_report", + mcp.WithDescription(RemoveBoardReportDescription), + mcp.WithDestructiveHintAnnotation(true), + mcp.WithString("board", mcp.Required(), mcp.Description(BoardIdDescription)), + mcp.WithString("reportId", mcp.Required(), mcp.Description("ID of the report to remove (from get_board).")), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + boardRef, err := RequiredParam[string](request, "board") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + reportID, err := RequiredParam[string](request, "reportId") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + + result, _, errResult := writeBoard(ctx, client, boardRef, "remove report", func(board *boards.Board) (UpdateBoardRequest, *mcp.CallToolResult) { + if _, ok := board.FindReport(reportID); !ok { + return UpdateBoardRequest{}, reportNotFound(board, reportID) + } + // The Control Plane leaves the layout alone when a report is + // deleted, so send the layout without the report's cell alongside + // the delete - derived from the board as it is now. + layout, err := boards.ParseLayout(board.Layout) + if err != nil { + return UpdateBoardRequest{}, mcp.NewToolResultError(err.Error()) + } + layoutJSON, _ := json.Marshal(layout.Without(reportID)) + return UpdateBoardRequest{ + Description: &board.Description, + Layout: layoutJSON, + Content: &BoardContentPatch{Action: "delete", ContentKind: "report", ContentID: reportID}, + }, nil + }) + if errResult != nil { + return errResult, nil + } + formatted, err := formatters.FormatBoard(result) + if err != nil { + return mcp.NewToolResultError(fmt.Sprintf("Failed to format board: %v", err)), nil + } + return mcp.NewToolResultText(formatted), nil + } + + return tool, handler +} + +// DeleteBoard creates a tool for deleting a board. +func DeleteBoard(client BoardRemover) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("delete_board", + mcp.WithDescription(DeleteBoardDescription), + mcp.WithDestructiveHintAnnotation(true), + mcp.WithString("board", mcp.Required(), mcp.Description(BoardIdDescription)), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + boardRef, err := RequiredParam[string](request, "board") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + board, errResult := fetchBoard(ctx, client, boardRef) + if errResult != nil { + return errResult, nil + } + if err := client.DeleteBoard(ctx, board.ID); err != nil { + return boardError("delete board", err), nil + } + return mcp.NewToolResultText(fmt.Sprintf("Deleted board %q (id: %s, slug: %s) and its %d report(s).", board.Name, board.ID, board.Slug, len(board.Content.Reports))), nil + } + + return tool, handler +} + +// renderedReport is one report's entry in the render_board result. +type renderedReport struct { + ID string `json:"id"` + Name string `json:"name,omitempty"` + Kind string `json:"kind"` + Query map[string]any `json:"query,omitempty"` + Data any `json:"data,omitempty"` + Error string `json:"error,omitempty"` +} + +// RenderBoard creates a tool that runs a board's reports and returns their data. +func RenderBoard(client BoardRenderer) (tool mcp.Tool, handler server.ToolHandlerFunc) { + tool = mcp.NewTool("render_board", + mcp.WithDescription(RenderBoardDescription), + mcp.WithReadOnlyHintAnnotation(true), + mcp.WithString("board", mcp.Required(), mcp.Description(BoardIdDescription)), + mcp.WithString("reportId", mcp.Description("Render only this report (from get_board). Renders every report when omitted.")), + mcp.WithString("scope", mcp.Description(BoardRenderScopeDescription), mcp.Enum("board", "environment")), + mcp.WithString("timeZone", mcp.Description(BoardRenderTimeZoneDescription)), + mcp.WithNumber("maxSamples", mcp.Description(InsightMaxSamplesDescription)), + ) + + handler = func(ctx context.Context, request mcp.CallToolRequest) (*mcp.CallToolResult, error) { + boardRef, err := RequiredParam[string](request, "board") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + reportID, err := OptionalParam[string](request, "reportId") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + scope, err := OptionalParam[string](request, "scope") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + if scope == "" { + scope = "board" + } + if scope != "board" && scope != "environment" { + return mcp.NewToolResultError("scope must be one of board, environment"), nil + } + maxSamples, err := OptionalIntParam(request, "maxSamples") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + maxSamples = min(max(maxSamples, 0), 500) + timeZone, err := OptionalParam[string](request, "timeZone") + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + loc, err := boards.ParseTimeZone(timeZone) + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + // One instant for every report, so they all cover the same range. + opts := boards.QueryOptions{Now: time.Now(), CurrentEnvironment: scope == "environment", Location: loc} + + board, errResult := fetchBoard(ctx, client, boardRef) + if errResult != nil { + return errResult, nil + } + reports, unplaced, err := board.OrderedReports() + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + if reportID != "" { + // A report asked for by ID is rendered even if the layout leaves + // it out. + r, ok := board.FindReport(reportID) + if !ok { + return mcp.NewToolResultError(fmt.Sprintf("report %q not found on board %q (reports: %s)", reportID, board.Name, strings.Join(board.ReportIDs(), ", "))), nil + } + reports = []boards.Report{*r} + } else { + // Render what the dashboard shows. Reports the layout leaves out + // are listed under unplaced, not queried. + reports = reports[:len(reports)-len(unplaced)] + } + + results := make([]renderedReport, len(reports)) + // The clients record each call in the context's DebugInfo, which is + // not safe for concurrent use. When debugging is on, every report + // query gets its own, merged into the call's once all are done. + debug := mcpcontext.GetDebugInfo(ctx) + reportDebug := make([]*mcpcontext.DebugInfo, len(reports)) + sem := make(chan struct{}, renderConcurrency) + var wg sync.WaitGroup + for i, r := range reports { + reportCtx := ctx + if debug != nil { + reportCtx, reportDebug[i] = mcpcontext.WithDebugInfo(ctx) + } + wg.Add(1) + go func() { + defer wg.Done() + sem <- struct{}{} + defer func() { <-sem }() + results[i] = renderReport(reportCtx, client, r, opts, maxSamples) + }() + } + wg.Wait() + if debug != nil { + perReport := make(map[string]*mcpcontext.DebugInfo, len(reports)) + for i, r := range reports { + perReport[r.ID] = reportDebug[i] + } + debug.Data["reports"] = perReport + } + + out := struct { + Board string `json:"board"` + Name string `json:"name"` + Scope string `json:"scope"` + TimeZone string `json:"timeZone"` + Reports []renderedReport `json:"reports"` + Unplaced []string `json:"unplaced,omitempty"` + }{Board: board.ID, Name: board.Name, Scope: scope, TimeZone: loc.String(), Reports: results} + if reportID == "" { + out.Unplaced = unplaced + } + formatted, err := formatters.FormatJSON(out) + if err != nil { + return mcp.NewToolResultError(fmt.Sprintf("Failed to format board data: %v", err)), nil + } + return mcp.NewToolResultText(formatted), nil + } + + return tool, handler +} + +func renderReport(ctx context.Context, client BoardInsightQuerier, r boards.Report, opts boards.QueryOptions, maxSamples int) renderedReport { + out := renderedReport{ID: r.ID, Name: r.Name, Kind: r.Kind} + q, err := boards.BuildQuery(r.Kind, r.Params, opts) + if err != nil { + out.Error = err.Error() + return out + } + out.Query = map[string]any{"endpoint": string(q.Endpoint)} + for k, v := range q.QueryParams() { + out.Query[k] = v + } + if q.CurrentEnvironment { + out.Query["env"] = "(current environment)" + } else if q.Env == "" { + out.Query["env"] = "(all environments)" + } + + raw, err := client.QueryBoardInsights(ctx, q) + if err != nil { + out.Error = insightErrorMessage(err) + return out + } + measure, _ := r.Params["measure"].(string) + data, err := formatters.ReportData(r.Kind, measure, raw, maxSamples) + if err != nil { + out.Error = fmt.Sprintf("failed to read report data: %v", err) + return out + } + out.Data = data + return out +} + +// fetchBoard loads a board for a tool that changes it. Every write reads the +// board first: to resend its description, to check a report exists, and to +// address it by ID even when the caller passed a slug. +func fetchBoard(ctx context.Context, client BoardGetter, ref string) (*boards.Board, *mcp.CallToolResult) { + raw, err := client.GetBoard(ctx, ref) + if err != nil { + return nil, boardError("get board", err) + } + board, err := boards.ParseBoard(raw) + if err != nil { + return nil, mcp.NewToolResultError(err.Error()) + } + return board, nil +} + +// writeBoard reads the board, builds an update from it, and sends it +// conditioned on the version it read. build must derive everything it takes +// from the board it is given. When a concurrent edit wins the race, the board +// is read again and the update rebuilt from it, so the caller's change is +// reapplied on top of the newer board instead of overwriting it. It returns +// the updated board and the board the update was built from. +func writeBoard(ctx context.Context, client BoardEditor, ref, action string, + build func(*boards.Board) (UpdateBoardRequest, *mcp.CallToolResult)) (string, *boards.Board, *mcp.CallToolResult) { + for attempt := 1; ; attempt++ { + board, errResult := fetchBoard(ctx, client, ref) + if errResult != nil { + return "", nil, errResult + } + req, errResult := build(board) + if errResult != nil { + return "", nil, errResult + } + // A Control Plane that predates versions returns none; then the write + // is unconditional, as before. + req.ExpectedVersion = board.Version + + result, err := client.UpdateBoard(ctx, board.ID, req) + if err == nil { + return result, board, nil + } + if !errors.Is(err, ErrBoardChanged) || attempt == boardWriteAttempts { + return "", nil, boardError(action, err) + } + // Read the same board again by the ID the first read resolved, not by + // what the caller passed: the concurrent change may have been to the + // slug, which could now be missing or belong to another board. + ref = board.ID + } +} + +func reportNotFound(board *boards.Board, reportID string) *mcp.CallToolResult { + return mcp.NewToolResultError(fmt.Sprintf("report %q not found on board %q (reports: %s)", reportID, board.Name, strings.Join(board.ReportIDs(), ", "))) +} + +// boardChangeResult formats a board returned by an update, together with the +// report the update was about. +func boardChangeResult(raw string, reportOf func(*boards.Board) string) (*mcp.CallToolResult, error) { + board, err := boards.ParseBoard(raw) + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + view, err := formatters.BoardView(board) + if err != nil { + return mcp.NewToolResultError(err.Error()), nil + } + out := struct { + ReportID string `json:"reportId,omitempty"` + Board formatters.FormattedBoard `json:"board"` + }{ReportID: reportOf(board), Board: view} + formatted, err := formatters.FormatJSON(out) + if err != nil { + return mcp.NewToolResultError(fmt.Sprintf("Failed to format board: %v", err)), nil + } + return mcp.NewToolResultText(formatted), nil +} + +// reportInput reads the params and filters arguments of a report tool. Params +// the kind does not understand are rejected. +func reportInput(request mcp.CallToolRequest, kind string) (map[string]any, map[string][]string, error) { + if !boards.IsKind(kind) { + return nil, nil, fmt.Errorf("unknown report kind %q (allowed: %s)", kind, strings.Join(boards.Kinds, ", ")) + } + var params map[string]any + switch v := request.GetArguments()["params"].(type) { + case nil: + case map[string]any: + params = v + case string: + if strings.TrimSpace(v) != "" { + if err := json.Unmarshal([]byte(v), ¶ms); err != nil { + return nil, nil, fmt.Errorf("params must be a JSON object: %w", err) + } + } + default: + return nil, nil, fmt.Errorf("params must be an object, got %T", v) + } + if err := boards.CheckParamKeys(kind, params); err != nil { + return nil, nil, err + } + + var filters map[string][]string + switch v := request.GetArguments()["filters"].(type) { + case nil: + case map[string]any: + filters = make(map[string][]string, len(v)) + for key, value := range v { + switch values := value.(type) { + case nil: + filters[key] = nil + case string: + filters[key] = []string{values} + case []any: + for _, item := range values { + s, ok := item.(string) + if !ok { + return nil, nil, fmt.Errorf("filters.%s must be a list of strings", key) + } + filters[key] = append(filters[key], s) + } + if filters[key] == nil { + filters[key] = []string{} + } + default: + return nil, nil, fmt.Errorf("filters.%s must be a list of strings", key) + } + } + default: + return nil, nil, fmt.Errorf("filters must be an object, got %T", v) + } + return params, filters, nil +} + +func parseLayoutParam(raw any) (boards.Layout, error) { + data, err := json.Marshal(raw) + if err != nil { + return boards.Layout{}, fmt.Errorf("layout is malformed: %w", err) + } + if s, ok := raw.(string); ok { + data = []byte(s) + } + var layout boards.Layout + if err := json.Unmarshal(data, &layout); err != nil { + return boards.Layout{}, fmt.Errorf("layout must be {\"version\": 1, \"rows\": [{\"cells\": [{\"id\": \"\"}]}]}: %w", err) + } + return layout, nil +} + +// boardError turns a client error into a tool error, replacing the Control +// Plane's refusal of API tokens with an actionable message. +func boardError(action string, err error) *mcp.CallToolResult { + msg := err.Error() + switch { + case errors.Is(err, ErrBoardsRequireUser), strings.Contains(msg, "API tokens are not supported"): + return mcp.NewToolResultError(ErrBoardsRequireUser.Error()) + case errors.Is(err, ErrBoardChanged): + return mcp.NewToolResultError(fmt.Sprintf("Failed to %s: the board kept changing while the change was applied (%d attempts), so nothing was written. Someone may be editing it; try again.", action, boardWriteAttempts)) + case strings.Contains(msg, "status 404"): + return mcp.NewToolResultError(fmt.Sprintf("Failed to %s: board not found (it may be private to another user). Use list_boards to find it. (%v)", action, err)) + case strings.Contains(msg, "status 500"): + // Older Control Planes answer an API-token caller with a bare 500. + return mcp.NewToolResultError(fmt.Sprintf("Failed to %s: %v. If this session uses an API token, note that boards require a signed-in user session (`testkube login`).", action, err)) + } + return mcp.NewToolResultError(fmt.Sprintf("Failed to %s: %v", action, err)) +} + +func insightErrorMessage(err error) string { + if strings.Contains(err.Error(), "status 403") { + return fmt.Sprintf("access denied - Insights may not be enabled for this organization (%v)", err) + } + return err.Error() +} diff --git a/pkg/mcp/tools/boards_test.go b/pkg/mcp/tools/boards_test.go new file mode 100644 index 00000000000..1b0b2902072 --- /dev/null +++ b/pkg/mcp/tools/boards_test.go @@ -0,0 +1,756 @@ +package tools + +import ( + "context" + "encoding/json" + "errors" + "fmt" + "strings" + "sync" + "testing" + "time" + + "github.com/mark3labs/mcp-go/mcp" + "github.com/mark3labs/mcp-go/server" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/kubeshop/testkube/pkg/mcp/boards" + mcpcontext "github.com/kubeshop/testkube/pkg/mcp/context" +) + +const testBoard = `{ + "id": "tkcbrd_1", "slug": "quality", "name": "Quality", "description": "Keep me", "shared": true, + "layout": {"version": 1, "rows": [{"id": "r1", "cells": [{"id": "aa"}, {"id": "bb"}]}, {"id": "r2", "cells": [{"id": "cc"}]}]}, + "content": {"reports": [ + {"id": "aa", "kind": "pass-fail", "name": "P/F", "params": {"measure": "ratio", "duration": "month", "filter": []}}, + {"id": "bb", "kind": "executions", "name": "Exec", "params": {"groupBy": "workflow", "measure": "count", "filter": [ + {"filterConfigurationKey": "workflow", "id": "w", "operator": "is", "value": ["api"]}]}}, + {"id": "cc", "kind": "time-series", "name": "TS", "params": {"measure": "", "aggregate": "sum", "filter": []}} + ]} +}` + +// fakeBoardClient records the requests the board tools make. +type fakeBoardClient struct { + mu sync.Mutex + + board string + getErr error + updateErr error + // updated is returned by UpdateBoard; defaults to board. + updated string + + slugAvailable bool + listParams ListBoardsParams + created *CreateBoardParams + updates []UpdateBoardRequest + updatedRef string + deleted string + queries []boards.InsightQuery + queryResult map[boards.Endpoint]string + queryErr map[boards.Endpoint]error + + // boardReads are the boards later reads return, in order. + boardReads []string + // updateErrs are the errors successive updates return, in order. + updateErrs []error + // boardsByRef, when set, serves boards by the ID or slug they are read by. + boardsByRef map[string]string + // readRefs records the ID or slug of every read. + readRefs []string + // queryDebug records the DebugInfo each insight query saw. + queryDebug []*mcpcontext.DebugInfo + // onUpdate, when set, runs as each update arrives, to change the board under it. + onUpdate func() +} + +func (f *fakeBoardClient) ListBoards(_ context.Context, p ListBoardsParams) (string, error) { + f.listParams = p + return `[{"id":"tkcbrd_1","slug":"quality","name":"Quality","shared":true,"analysisCount":3}]`, nil +} + +func (f *fakeBoardClient) GetBoard(_ context.Context, ref string) (string, error) { + f.readRefs = append(f.readRefs, ref) + if f.boardsByRef != nil { + if b, ok := f.boardsByRef[ref]; ok { + return b, nil + } + return "", errors.New("API returned status 404: board not found") + } + // Each read after the first returns the next board of boardReads, as if + // someone edited the board in between. + if len(f.boardReads) > 0 { + f.board, f.boardReads = f.boardReads[0], f.boardReads[1:] + } + return f.board, f.getErr +} + +func (f *fakeBoardClient) CheckBoardSlug(context.Context, string) (bool, error) { + return f.slugAvailable, nil +} + +func (f *fakeBoardClient) CreateBoard(_ context.Context, p CreateBoardParams) (string, error) { + f.created = &p + return f.board, nil +} + +func (f *fakeBoardClient) UpdateBoard(_ context.Context, ref string, r UpdateBoardRequest) (string, error) { + if f.onUpdate != nil { + f.onUpdate() + } + f.updatedRef = ref + f.updates = append(f.updates, r) + if len(f.updateErrs) > 0 { + err := f.updateErrs[0] + f.updateErrs = f.updateErrs[1:] + if err != nil { + return "", err + } + } + if f.updated != "" { + return f.updated, f.updateErr + } + return f.board, f.updateErr +} + +func (f *fakeBoardClient) DeleteBoard(_ context.Context, ref string) error { + f.deleted = ref + return nil +} + +func (f *fakeBoardClient) QueryBoardInsights(ctx context.Context, q boards.InsightQuery) (string, error) { + f.mu.Lock() + f.queries = append(f.queries, q) + f.queryDebug = append(f.queryDebug, mcpcontext.GetDebugInfo(ctx)) + f.mu.Unlock() + // Record the call in the context's DebugInfo without a lock, as the real + // clients do. + if d := mcpcontext.GetDebugInfo(ctx); d != nil { + d.Source = "fake" + for i := 0; i < 50; i++ { + d.Data[fmt.Sprintf("k%d", i)] = i + } + d.Data["url"] = string(q.Endpoint) + } + if err := f.queryErr[q.Endpoint]; err != nil { + return "", err + } + return f.queryResult[q.Endpoint], nil +} + +func callTool(t *testing.T, handler func(context.Context, mcp.CallToolRequest) (*mcp.CallToolResult, error), args map[string]any) *mcp.CallToolResult { + t.Helper() + request := mcp.CallToolRequest{} + request.Params.Arguments = args + result, err := handler(context.Background(), request) + require.NoError(t, err) + require.NotNil(t, result) + return result +} + +func (f *fakeBoardClient) lastUpdate(t *testing.T) UpdateBoardRequest { + t.Helper() + require.NotEmpty(t, f.updates, "expected an update request") + return f.updates[len(f.updates)-1] +} + +func TestListBoards(t *testing.T) { + f := &fakeBoardClient{} + tool, handler := ListBoards(f) + assert.Equal(t, "list_boards", tool.Name) + require.NotNil(t, tool.Annotations.ReadOnlyHint) + assert.True(t, *tool.Annotations.ReadOnlyHint) + + result := callTool(t, handler, map[string]any{"visibility": "private", "favorite": "user", "pageSize": float64(1)}) + require.False(t, result.IsError, getResultText(result)) + assert.Equal(t, ListBoardsParams{Private: true, UserFavorite: true, PageSize: 1}, f.listParams) + assert.Contains(t, getResultText(result), `"hasMore":true`) + assert.Contains(t, getResultText(result), `"reports":3`) + + result = callTool(t, handler, map[string]any{"visibility": "team"}) + assert.True(t, result.IsError) +} + +func TestGetBoard_ReportsInLayoutOrder(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := GetBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality"}) + require.False(t, result.IsError, getResultText(result)) + text := getResultText(result) + assert.Contains(t, text, `"layout":[["aa","bb"],["cc"]]`) + assert.Less(t, strings.Index(text, `"id":"aa"`), strings.Index(text, `"id":"cc"`)) +} + +func TestBoardTools_TokenRejection(t *testing.T) { + f := &fakeBoardClient{getErr: ErrBoardsRequireUser} + _, handler := GetBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality"}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), "testkube login") + + // The Control Plane's own refusal, as a newer version reports it. + f.getErr = errors.New("API returned status 403: boards with API tokens are not supported") + result = callTool(t, handler, map[string]any{"board": "quality"}) + assert.Contains(t, getResultText(result), "testkube login") +} + +func TestCreateBoard(t *testing.T) { + t.Run("refuses a taken slug", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard, slugAvailable: false} + _, handler := CreateBoard(f) + result := callTool(t, handler, map[string]any{"name": "Quality", "slug": "quality"}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), "already used") + assert.Nil(t, f.created) + }) + + t.Run("creates", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard, slugAvailable: true} + _, handler := CreateBoard(f) + result := callTool(t, handler, map[string]any{"name": "Quality", "slug": "quality", "private": true, "description": "d"}) + require.False(t, result.IsError, getResultText(result)) + assert.Equal(t, &CreateBoardParams{Name: "Quality", Slug: "quality", Description: "d", IsPrivate: true}, f.created) + }) +} + +func TestUpdateBoard(t *testing.T) { + t.Run("keeps the description when only the name changes", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "name": "Renamed"}) + require.False(t, result.IsError, getResultText(result)) + u := f.lastUpdate(t) + require.NotNil(t, u.Name) + assert.Equal(t, "Renamed", *u.Name) + require.NotNil(t, u.Description) + assert.Equal(t, "Keep me", *u.Description) + assert.Equal(t, "tkcbrd_1", f.updatedRef, "the update addresses the board by ID even when given a slug") + }) + + t.Run("clears the description on request", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoard(f) + callTool(t, handler, map[string]any{"board": "quality", "description": ""}) + assert.Equal(t, "", *f.lastUpdate(t).Description) + }) + + t.Run("validates the layout", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "layout": map[string]any{ + "version": float64(1), "rows": []any{map[string]any{"cells": []any{map[string]any{"id": "aa"}}}}, + }}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), "leaves out reports bb, cc") + assert.Empty(t, f.updates) + + result = callTool(t, handler, map[string]any{"board": "quality", "layout": map[string]any{ + "version": float64(1), "rows": []any{ + map[string]any{"cells": []any{map[string]any{"id": "cc"}}}, + map[string]any{"cells": []any{map[string]any{"id": "aa"}, map[string]any{"id": "bb"}}}, + }, + }}) + require.False(t, result.IsError, getResultText(result)) + var layout boards.Layout + require.NoError(t, json.Unmarshal(f.lastUpdate(t).Layout, &layout)) + assert.Equal(t, [][]string{{"cc"}, {"aa", "bb"}}, layout.RowIDs()) + }) + + t.Run("refuses an empty update", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality"}) + assert.True(t, result.IsError) + assert.Empty(t, f.updates) + }) +} + +func TestAddBoardReport(t *testing.T) { + t.Run("normalizes params, applies filters and returns the new report ID", func(t *testing.T) { + updated := strings.Replace(testBoard, `{"id": "aa",`, `{"id": "dd", "kind": "workflows", "name": "New"}, {"id": "aa",`, 1) + f := &fakeBoardClient{board: testBoard, updated: updated} + _, handler := AddBoardReport(f) + result := callTool(t, handler, map[string]any{ + "board": "quality", "kind": "time-series", "name": "Duration", + "params": map[string]any{"measure": "execution-duration", "aggregate": "avg"}, + "filters": map[string]any{"workflow": []any{"api"}, "environment": "tkcenv_1"}, + }) + require.False(t, result.IsError, getResultText(result)) + assert.Contains(t, getResultText(result), `"reportId":"dd"`) + + u := f.lastUpdate(t) + assert.Equal(t, "Keep me", *u.Description) + assert.Nil(t, u.Layout, "the Control Plane places a new report itself") + require.NotNil(t, u.Content) + assert.Equal(t, "create", u.Content.Action) + assert.Equal(t, "report", u.Content.ContentKind) + p := u.Content.ContentData.Params + assert.Equal(t, "execution-duration", p["measure"]) + assert.Equal(t, "avg", p["aggregate"]) + assert.Equal(t, "week", p["duration"]) + assert.Equal(t, "bar", p["chartType"]) + filters, err := boards.ParseFilters(p["filter"]) + require.NoError(t, err) + assert.Equal(t, []string{"api"}, boards.FilterValues(filters, boards.FilterWorkflow)) + assert.Equal(t, []string{"tkcenv_1"}, boards.FilterValues(filters, boards.FilterEnvironment)) + }) + + t.Run("rejects unknown params before touching the board", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := AddBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "kind": "workflows", "name": "W", "params": map[string]any{"measure": "count"}}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), "unknown params") + assert.Empty(t, f.updates) + }) + + t.Run("accepts params as a JSON string", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := AddBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "kind": "pass-fail", "name": "P", "params": `{"measure":"failed-count"}`}) + require.False(t, result.IsError, getResultText(result)) + assert.Equal(t, "failed-count", f.lastUpdate(t).Content.ContentData.Params["measure"]) + }) +} + +func TestUpdateBoardReport(t *testing.T) { + t.Run("rejects an unknown report", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "zz", "name": "x"}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), `report "zz" not found`) + assert.Empty(t, f.updates) + }) + + t.Run("merges params and keeps the rest of the report", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "bb", "params": map[string]any{"measure": "duration"}}) + require.False(t, result.IsError, getResultText(result)) + u := f.lastUpdate(t) + assert.Equal(t, "Keep me", *u.Description) + assert.Equal(t, "update", u.Content.Action) + assert.Equal(t, "bb", u.Content.ContentID) + d := u.Content.ContentData + assert.Equal(t, "executions", d.Kind) + assert.Equal(t, "Exec", d.Name) + assert.Equal(t, "duration", d.Params["measure"]) + assert.Equal(t, "workflow", d.Params["groupBy"], "unchanged params survive a merge") + filters, _ := boards.ParseFilters(d.Params["filter"]) + assert.Equal(t, []string{"api"}, boards.FilterValues(filters, boards.FilterWorkflow)) + }) + + t.Run("replaceParams starts from defaults", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoardReport(f) + callTool(t, handler, map[string]any{"board": "quality", "reportId": "bb", "replaceParams": true, "params": map[string]any{"measure": "duration"}}) + assert.Equal(t, "status", f.lastUpdate(t).Content.ContentData.Params["groupBy"]) + }) + + t.Run("changing the kind drops the old params", func(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "bb", "kind": "workflows"}) + require.False(t, result.IsError, getResultText(result)) + d := f.lastUpdate(t).Content.ContentData + assert.Equal(t, "workflows", d.Kind) + assert.NotContains(t, d.Params, "groupBy") + }) +} + +func TestRemoveBoardReport(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + tool, handler := RemoveBoardReport(f) + require.NotNil(t, tool.Annotations.DestructiveHint) + assert.True(t, *tool.Annotations.DestructiveHint) + + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "cc"}) + require.False(t, result.IsError, getResultText(result)) + u := f.lastUpdate(t) + assert.Equal(t, "delete", u.Content.Action) + assert.Equal(t, "cc", u.Content.ContentID) + assert.Equal(t, "Keep me", *u.Description) + var layout boards.Layout + require.NoError(t, json.Unmarshal(u.Layout, &layout)) + assert.Equal(t, [][]string{{"aa", "bb"}}, layout.RowIDs(), "the cell and its now-empty row are dropped") + + result = callTool(t, handler, map[string]any{"board": "quality", "reportId": "zz"}) + assert.True(t, result.IsError) +} + +func TestDeleteBoard(t *testing.T) { + f := &fakeBoardClient{board: testBoard} + tool, handler := DeleteBoard(f) + assert.True(t, *tool.Annotations.DestructiveHint) + + result := callTool(t, handler, map[string]any{"board": "quality"}) + require.False(t, result.IsError, getResultText(result)) + assert.Equal(t, "tkcbrd_1", f.deleted) + assert.Contains(t, getResultText(result), `Deleted board "Quality"`) + assert.Contains(t, getResultText(result), "3 report(s)") +} + +func TestRenderBoard(t *testing.T) { + newClient := func() *fakeBoardClient { + return &fakeBoardClient{ + board: testBoard, + queryResult: map[boards.Endpoint]string{ + boards.EndpointStats: `{"ratioStats":{"total":90,"values":[["2026-09-01",80],["2026-09-02",100]]},"totalStats":{"total":10,"values":[]},"failedStats":{"total":1,"values":[]}}`, + boards.EndpointExecutions: `{"count":{"total":5,"values":[["api",5]]},"duration":{"total":0,"values":[]}}`, + }, + queryErr: map[boards.Endpoint]error{}, + } + } + + t.Run("renders every report and isolates failures", func(t *testing.T) { + f := newClient() + f.queryErr[boards.EndpointExecutions] = fmt.Errorf("API returned status 403: forbidden") + _, handler := RenderBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality"}) + require.False(t, result.IsError, getResultText(result)) + + var out struct { + Scope string `json:"scope"` + Reports []struct { + ID string `json:"id"` + Query map[string]any `json:"query"` + Data map[string]any `json:"data"` + Error string `json:"error"` + } `json:"reports"` + } + require.NoError(t, json.Unmarshal([]byte(getResultText(result)), &out)) + assert.Equal(t, "board", out.Scope) + require.Len(t, out.Reports, 3) + assert.Equal(t, []string{"aa", "bb", "cc"}, []string{out.Reports[0].ID, out.Reports[1].ID, out.Reports[2].ID}) + + assert.Equal(t, float64(90), out.Reports[0].Data["total"]) + assert.Equal(t, "(all environments)", out.Reports[0].Query["env"]) + assert.Contains(t, out.Reports[1].Error, "Insights may not be enabled") + assert.Contains(t, out.Reports[2].Error, "no measure", "a time-series report without a measure is not queried") + assert.Len(t, f.queries, 2) + }) + + t.Run("environment scope asks for the current environment", func(t *testing.T) { + f := newClient() + _, handler := RenderBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "aa", "scope": "environment"}) + require.False(t, result.IsError, getResultText(result)) + require.Len(t, f.queries, 1) + assert.True(t, f.queries[0].CurrentEnvironment) + assert.Contains(t, getResultText(result), "(current environment)") + }) +} + +func TestRenderBoard_TimeZone(t *testing.T) { + f := &fakeBoardClient{ + board: testBoard, + queryResult: map[boards.Endpoint]string{boards.EndpointStats: `{"ratioStats":{"total":1,"values":[]},"totalStats":{"total":0,"values":[]},"failedStats":{"total":0,"values":[]}}`}, + queryErr: map[boards.Endpoint]error{}, + } + _, handler := RenderBoard(f) + + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "aa", "timeZone": "Asia/Tokyo"}) + require.False(t, result.IsError, getResultText(result)) + assert.Contains(t, getResultText(result), `"timeZone":"Asia/Tokyo"`) + require.Len(t, f.queries, 1) + // A relative range ends at the start of tomorrow in Tokyo, which is 15:00 UTC. + end := f.queries[0].EndDate.In(time.FixedZone("JST", 9*60*60)) + assert.Equal(t, 0, end.Hour()) + assert.Equal(t, 15, end.UTC().Hour()) + + result = callTool(t, handler, map[string]any{"board": "quality", "timeZone": "Nowhere/Land"}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), "IANA time zone") +} + +// boardAt returns the test board as read at updatedAt, with a description. +func boardAt(version int64, description string) string { + b := strings.Replace(testBoard, `"shared": true,`, fmt.Sprintf(`"shared": true, "version": %d,`, version), 1) + return strings.Replace(b, `"description": "Keep me"`, `"description": "`+description+`"`, 1) +} + +func TestBoardWrites_SendTheVersionTheyRead(t *testing.T) { + const at = int64(7) + tests := []struct { + name string + tool func(BoardEditor) (mcp.Tool, server.ToolHandlerFunc) + args map[string]any + }{ + {"update_board", UpdateBoard, map[string]any{"board": "quality", "name": "Renamed"}}, + {"add_board_report", AddBoardReport, map[string]any{"board": "quality", "kind": "workflows", "name": "W"}}, + {"update_board_report", UpdateBoardReport, map[string]any{"board": "quality", "reportId": "bb", "name": "Renamed"}}, + {"remove_board_report", RemoveBoardReport, map[string]any{"board": "quality", "reportId": "cc"}}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + f := &fakeBoardClient{board: boardAt(at, "Keep me")} + _, handler := tt.tool(f) + result := callTool(t, handler, tt.args) + require.False(t, result.IsError, getResultText(result)) + require.NotNil(t, f.lastUpdate(t).ExpectedVersion) + assert.Equal(t, at, *f.lastUpdate(t).ExpectedVersion) + }) + } +} + +func TestBoardWrites_RebuildFromTheNewerBoardAfterAConflict(t *testing.T) { + const first, second = int64(3), int64(4) + changed := fmt.Errorf("%w: API returned status 409", ErrBoardChanged) + + t.Run("a concurrent description edit is kept", func(t *testing.T) { + f := &fakeBoardClient{ + boardReads: []string{boardAt(first, "old"), boardAt(second, "edited meanwhile")}, + updateErrs: []error{changed, nil}, + } + _, handler := UpdateBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "name": "Renamed"}) + require.False(t, result.IsError, getResultText(result)) + + require.Len(t, f.updates, 2) + retry := f.updates[1] + assert.Equal(t, second, *retry.ExpectedVersion) + assert.Equal(t, "edited meanwhile", *retry.Description, "the stale description must not be written back") + assert.Equal(t, "Renamed", *retry.Name) + }) + + t.Run("a report removal is recomputed from the newer layout", func(t *testing.T) { + newer := strings.Replace(boardAt(second, "Keep me"), + `{"id": "r2", "cells": [{"id": "cc"}]}`, + `{"id": "r2", "cells": [{"id": "cc"}, {"id": "bb"}]}`, 1) + newer = strings.Replace(newer, `[{"id": "aa"}, {"id": "bb"}]`, `[{"id": "aa"}]`, 1) + f := &fakeBoardClient{ + boardReads: []string{boardAt(first, "Keep me"), newer}, + updateErrs: []error{changed, nil}, + } + _, handler := RemoveBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "cc"}) + require.False(t, result.IsError, getResultText(result)) + + var layout boards.Layout + require.NoError(t, json.Unmarshal(f.lastUpdate(t).Layout, &layout)) + assert.Equal(t, [][]string{{"aa"}, {"bb"}}, layout.RowIDs(), "the move of bb made meanwhile must survive the removal") + }) + + t.Run("a report edit merges into the newer params", func(t *testing.T) { + newer := strings.Replace(boardAt(second, "Keep me"), `"groupBy": "workflow"`, `"groupBy": "status"`, 1) + f := &fakeBoardClient{ + boardReads: []string{boardAt(first, "Keep me"), newer}, + updateErrs: []error{changed, nil}, + } + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "bb", "params": map[string]any{"measure": "duration"}}) + require.False(t, result.IsError, getResultText(result)) + + params := f.lastUpdate(t).Content.ContentData.Params + assert.Equal(t, "duration", params["measure"]) + assert.Equal(t, "status", params["groupBy"], "a param changed meanwhile and not named by the caller must be kept") + }) + + t.Run("a report removed meanwhile is reported, not recreated", func(t *testing.T) { + gone := strings.Replace(boardAt(second, "Keep me"), `{"id": "bb", "kind": "executions"`, `{"id": "zz", "kind": "executions"`, 1) + f := &fakeBoardClient{ + boardReads: []string{boardAt(first, "Keep me"), gone}, + updateErrs: []error{changed}, + } + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "bb", "name": "x"}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), `report "bb" not found`) + assert.Len(t, f.updates, 1) + }) + + t.Run("gives up after a bounded number of attempts", func(t *testing.T) { + f := &fakeBoardClient{ + board: boardAt(first, "Keep me"), + updateErrs: []error{changed, changed, changed, nil}, + } + _, handler := UpdateBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "name": "Renamed"}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), "kept changing") + assert.Len(t, f.updates, boardWriteAttempts) + }) + + t.Run("other errors are not retried", func(t *testing.T) { + f := &fakeBoardClient{ + board: boardAt(first, "Keep me"), + updateErrs: []error{errors.New("API returned status 403: forbidden")}, + } + _, handler := UpdateBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "name": "Renamed"}) + assert.True(t, result.IsError) + assert.Len(t, f.updates, 1) + }) +} + +func TestBoardWrites_AreUnconditionalWithoutAVersion(t *testing.T) { + // A Control Plane that predates versions returns none, and would ignore + // an expectation anyway, so none is sent. + f := &fakeBoardClient{board: testBoard} + _, handler := UpdateBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "name": "Renamed"}) + require.False(t, result.IsError, getResultText(result)) + assert.Nil(t, f.lastUpdate(t).ExpectedVersion) +} + +func TestBoardWrites_RetryReadsTheSameBoardByID(t *testing.T) { + // The caller names the board by slug. Before the write lands, someone + // renames its slug and another board takes the old one over. + original := boardAt(3, "Keep me") + renamed := strings.Replace(boardAt(4, "Keep me"), `"slug": "quality"`, `"slug": "quality-renamed"`, 1) + impostor := strings.Replace(strings.Replace(boardAt(9, "Someone else's"), `"id": "tkcbrd_1"`, `"id": "tkcbrd_2"`, 1), `"name": "Quality"`, `"name": "Other"`, 1) + + f := &fakeBoardClient{ + boardsByRef: map[string]string{"quality": original}, + updateErrs: []error{fmt.Errorf("%w: API returned status 409", ErrBoardChanged), nil}, + updated: renamed, + } + _, handler := UpdateBoard(f) + // While the first update is in flight, the slug moves to the other board. + first := true + f.onUpdate = func() { + if first { + f.boardsByRef = map[string]string{"quality": impostor, "tkcbrd_1": renamed, "quality-renamed": renamed} + first = false + } + } + + result := callTool(t, handler, map[string]any{"board": "quality", "name": "Renamed"}) + require.False(t, result.IsError, getResultText(result)) + + assert.Equal(t, []string{"quality", "tkcbrd_1"}, f.readRefs, "the retry must read the board the first read resolved, by ID") + assert.Equal(t, "tkcbrd_1", f.updatedRef) + require.Len(t, f.updates, 2) + assert.Equal(t, int64(4), *f.updates[1].ExpectedVersion, "the retry is built from the renamed board, not the one now holding the slug") +} + +func TestRenderBoard_GivesEachReportItsOwnDebugInfo(t *testing.T) { + // testBoard has three reports; aa and bb query in parallel (cc has no + // measure and is not queried). Rendering repeatedly makes a shared, + // unsynchronized DebugInfo likely to fail on concurrent map writes. + for run := 0; run < 20; run++ { + f := &fakeBoardClient{ + board: testBoard, + queryResult: map[boards.Endpoint]string{ + boards.EndpointStats: `{"ratioStats":{"total":1,"values":[]},"totalStats":{"total":0,"values":[]},"failedStats":{"total":0,"values":[]}}`, + boards.EndpointExecutions: `{"count":{"total":0,"values":[]},"duration":{"total":0,"values":[]}}`, + }, + queryErr: map[boards.Endpoint]error{}, + } + _, handler := RenderBoard(f) + ctx, debug := mcpcontext.WithDebugInfo(context.Background()) + request := mcp.CallToolRequest{} + request.Params.Arguments = map[string]any{"board": "quality"} + result, err := handler(ctx, request) + require.NoError(t, err) + require.False(t, result.IsError, getResultText(result)) + + require.Len(t, f.queryDebug, 2) + assert.NotSame(t, debug, f.queryDebug[0], "a report query must not write the call's shared DebugInfo") + assert.NotSame(t, f.queryDebug[0], f.queryDebug[1], "each report query gets its own DebugInfo") + + perReport, ok := debug.Data["reports"].(map[string]*mcpcontext.DebugInfo) + require.True(t, ok, "the per-report debug info is merged into the call's") + assert.Equal(t, string(boards.EndpointStats), perReport["aa"].Data["url"]) + assert.Equal(t, string(boards.EndpointExecutions), perReport["bb"].Data["url"]) + assert.Empty(t, perReport["cc"].Data, "a report that is not queried records nothing") + } +} + +func TestRenderBoard_WithoutDebugAddsNothing(t *testing.T) { + f := &fakeBoardClient{board: testBoard, queryResult: map[boards.Endpoint]string{}, queryErr: map[boards.Endpoint]error{}} + _, handler := RenderBoard(f) + callTool(t, handler, map[string]any{"board": "quality"}) + for _, d := range f.queryDebug { + assert.Nil(t, d, "no DebugInfo is created when debugging is off") + } +} + +func TestRenderBoard_SkipsReportsTheLayoutLeavesOut(t *testing.T) { + // cc is on the board but not in the layout, so the dashboard does not show it. + unplaced := strings.Replace(testBoard, `, {"id": "r2", "cells": [{"id": "cc"}]}`, ``, 1) + unplaced = strings.Replace(unplaced, `"measure": "", "aggregate": "sum"`, `"measure": "execution-count", "aggregate": "sum"`, 1) + newClient := func() *fakeBoardClient { + return &fakeBoardClient{ + board: unplaced, + queryResult: map[boards.Endpoint]string{ + boards.EndpointStats: `{"ratioStats":{"total":1,"values":[]},"totalStats":{"total":0,"values":[]},"failedStats":{"total":0,"values":[]}}`, + boards.EndpointExecutions: `{"count":{"total":0,"values":[]},"duration":{"total":0,"values":[]}}`, + boards.EndpointSeries: `[]`, + }, + queryErr: map[boards.Endpoint]error{}, + } + } + var out struct { + Reports []struct { + ID string `json:"id"` + } `json:"reports"` + Unplaced []string `json:"unplaced"` + } + + t.Run("the default render covers what the dashboard shows", func(t *testing.T) { + f := newClient() + _, handler := RenderBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality"}) + require.False(t, result.IsError, getResultText(result)) + require.NoError(t, json.Unmarshal([]byte(getResultText(result)), &out)) + + ids := []string{} + for _, r := range out.Reports { + ids = append(ids, r.ID) + } + assert.Equal(t, []string{"aa", "bb"}, ids) + assert.Equal(t, []string{"cc"}, out.Unplaced, "the left-out report is still named") + for _, q := range f.queries { + assert.NotEqual(t, boards.EndpointSeries, q.Endpoint, "the left-out report must not be queried") + } + }) + + t.Run("a left-out report asked for by ID is rendered", func(t *testing.T) { + f := newClient() + _, handler := RenderBoard(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "cc"}) + require.False(t, result.IsError, getResultText(result)) + require.NoError(t, json.Unmarshal([]byte(getResultText(result)), &out)) + require.Len(t, out.Reports, 1) + assert.Equal(t, "cc", out.Reports[0].ID) + require.Len(t, f.queries, 1) + assert.Equal(t, boards.EndpointSeries, f.queries[0].Endpoint) + }) +} + +func TestUpdateBoardReport_KeepsWhatTheReportShows(t *testing.T) { + // aa is a pass-fail report stored without a duration, so the dashboard + // renders a week. A new pass-fail report defaults to a month. + noDuration := strings.Replace(testBoard, `"measure": "ratio", "duration": "month", "filter": []`, `"measure": "ratio", "filter": []`, 1) + + t.Run("an unrelated edit adds no creation defaults", func(t *testing.T) { + f := &fakeBoardClient{board: noDuration} + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "aa", "params": map[string]any{"measure": "failed-count"}}) + require.False(t, result.IsError, getResultText(result)) + params := f.lastUpdate(t).Content.ContentData.Params + assert.Equal(t, "failed-count", params["measure"]) + assert.NotContains(t, params, "duration", "the report's window must not change") + }) + + t.Run("changing the kind starts from that kind's defaults", func(t *testing.T) { + f := &fakeBoardClient{board: noDuration} + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "aa", "kind": "workflows"}) + require.False(t, result.IsError, getResultText(result)) + assert.Equal(t, "month", f.lastUpdate(t).Content.ContentData.Params["duration"]) + }) + + t.Run("a bad label filter is refused, not stored", func(t *testing.T) { + f := &fakeBoardClient{board: noDuration} + _, handler := UpdateBoardReport(f) + result := callTool(t, handler, map[string]any{"board": "quality", "reportId": "aa", "params": map[string]any{ + "filter": []any{map[string]any{"filterConfigurationKey": "labels-v2", "operator": "is", "value": "team=core"}}, + }}) + assert.True(t, result.IsError) + assert.Contains(t, getResultText(result), "labels-v2") + assert.Empty(t, f.updates) + }) +} diff --git a/pkg/mcp/tools/descriptions.go b/pkg/mcp/tools/descriptions.go index a9fb014d8d8..ba626c80baf 100644 --- a/pkg/mcp/tools/descriptions.go +++ b/pkg/mcp/tools/descriptions.go @@ -134,3 +134,68 @@ Filter with the same 'measure'/identity/workflow/status/tag/date filters as get_ InsightTagFilterDescription = `Filter executions by tag. Tag values may contain commas, so multiple predicates must be passed as a JSON array, e.g. ["release=2026.07,hotfix","bug"]. ` + `A plain (non-JSON) string is treated as a single predicate verbatim. Each predicate supports key=value (exact), key=~pattern (regex), or a bare key (existence).` ) + +// Insights board tool descriptions +const ( + boardSessionNote = ` Boards belong to the organization, not to one environment. Requires a signed-in user session: API tokens cannot use boards.` + + ListBoardsDescription = `List the Insights boards visible to you: boards shared with the organization and your private boards. +A board is a saved dashboard of reports (charts) over execution data. Use get_board to see a board's reports and render_board to get their numbers.` + boardSessionNote + + GetBoardDescription = `Get an Insights board: its details, its reports (id, kind, name, params) in the order the dashboard shows them, and its layout as rows of report IDs. +Reports listed under 'unplaced' are on the board but not in its layout, so the dashboard does not show them.` + boardSessionNote + + CreateBoardDescription = `Create an empty Insights board. Shared with the organization by default; set private to keep it to yourself. +Add reports to it with add_board_report.` + boardSessionNote + + UpdateBoardDescription = `Change an Insights board's name, description, slug, visibility or layout. Fields you omit keep their current values. +To change the reports on a board, use add_board_report, update_board_report and remove_board_report instead.` + boardSessionNote + + AddBoardReportDescription = `Add a report (chart) to an Insights board. The report is placed on a new row at the bottom of the board, and its ID is returned. +Kinds: 'pass-fail' (pass/fail ratio or counts over time), 'executions' (execution count or average duration grouped by status, workflow, label or tag), +'workflows' (per-workflow execution, failure and duration summary), 'time-series' (any measure over time, optionally segmented - including performance and resource metrics). +Params you omit get the same defaults the dashboard gives a new report.` + boardSessionNote + + UpdateBoardReportDescription = `Change a report on an Insights board: its name, description, kind or params. Params are merged into the current ones unless replaceParams is true. +The report keeps its place on the board.` + boardSessionNote + + RemoveBoardReportDescription = `Remove a report from an Insights board. This cannot be undone. Only call it when the user asked to remove this report.` + boardSessionNote + + DeleteBoardDescription = `Permanently delete an Insights board and all of its reports. This cannot be undone. +Only call it when the user explicitly asked to delete this board. Deleting a shared board requires an organization admin; a private board can be deleted by its owner.` + boardSessionNote + + RenderBoardDescription = `Render an Insights board: run the query of each report the dashboard shows and return its numbers, together with the query and date range used. Reports the board's layout leaves out are listed under 'unplaced' and not rendered, unless one is asked for with reportId. +A report that fails is returned with an error without failing the others. Relative durations end at the start of tomorrow in timeZone, as the dashboard ends them at the viewer's local midnight - pass the user's time zone to get the numbers they see.` + boardSessionNote + + // Board tool parameter descriptions + BoardIdDescription = "The board's ID or slug (from list_boards)." + + BoardReportKindDescription = "Report kind: 'pass-fail', 'executions', 'workflows' or 'time-series'." + + BoardReportParamsDescription = `Report params (a JSON object). Common to every kind: +- duration: 'day', 'week', 'month' or 'quarter' (120 days) ending at the start of tomorrow; or set from/to (RFC3339) for a fixed range. +- filter: dashboard filter list; prefer the 'filters' argument instead. +Per kind: +- pass-fail: measure 'ratio' (default), 'failed-count' or 'total-count'. Default duration: month. +- executions: groupBy 'status' (default), 'workflow', a label key, or 'tag:'; measure 'count' (default) or 'duration' (average, ms). +- workflows: no extra params. Default duration: month. +- time-series: measure (default 'execution-count') - 'execution-count', 'execution-duration', 'case-count', + a resource measure '-', + or a granular metric key from list_insight_metric_keys; aggregate 'sum' (default), 'avg', 'min', 'max' or 'count'; + segment (break down by 'status', 'workflow', a label key, 'tag:' or an identity field; defaults to 'status' for execution-count, '' for none); + chartType 'bar' (default), 'bar-grouped', 'bar-normalized', 'line-stacked', 'line', 'area-normalized', 'heatmap' or 'horizon'; overlaySuccessRate (boolean). +Example: {"measure": "execution-duration", "aggregate": "avg", "segment": "workflow", "duration": "month", "chartType": "line"}` + + BoardReportFiltersDescription = `Filters narrowing the report, as {key: [values]}. Keys: 'environment' (environment IDs; omit for every environment), +'workflow' (workflow names), 'status' (e.g. 'passed', 'failed'), 'labels' (label selectors 'key=value'), 'tags' (tag predicates 'key=value'). +For time-series reports any other key is a granular identity filter (e.g. 'testcase'). Values prefixed with '~' match as a regex. +Example: {"workflow": ["api-tests"], "status": ["failed"], "environment": ["tkcenv_..."]}` + + BoardLayoutDescription = `Advanced: the board layout, to reorder reports or put several on one row. Format: {"version": 1, "rows": [{"cells": [{"id": ""}, ...]}, ...]}. +Every report on the board must appear exactly once. Get the current layout and report IDs with get_board.` + + BoardRenderTimeZoneDescription = `IANA time zone of the person viewing the board, e.g. 'Europe/Berlin' or 'America/New_York'. The dashboard ends a relative range (day, week, month, quarter) at the viewer's local midnight, so use the user's zone to match what they see. Default: UTC.` + + BoardRenderScopeDescription = `'board' (default) renders what the dashboard shows: each report's own environment filter, or every environment when it has none. +'environment' limits every report to the current environment instead.` +) diff --git a/test/env/dev/dev-postgres/workflows/gh-quality-loop.yaml b/test/env/dev/dev-postgres/workflows/gh-quality-loop.yaml index 33b402203cc..d0bdc8f3fb4 100644 --- a/test/env/dev/dev-postgres/workflows/gh-quality-loop.yaml +++ b/test/env/dev/dev-postgres/workflows/gh-quality-loop.yaml @@ -193,7 +193,7 @@ spec: workingDir: /data/repo resources: limits: - cpu: 4 + cpu: 8 memory: 8Gi requests: cpu: 2 @@ -213,6 +213,7 @@ spec: value: https://proxy.golang.org,direct shell: |- golangci-lint run -c ./.golangci.yml \ + --verbose \ --timeout=20m \ --issues-exit-code=1 \ --max-issues-per-linter=0 \