Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

roachtest: cdc/cloud-sink-gcs/rangefeed=true failed #87677

Closed
cockroach-teamcity opened this issue Sep 9, 2022 · 5 comments · Fixed by #87817
Closed

roachtest: cdc/cloud-sink-gcs/rangefeed=true failed #87677

cockroach-teamcity opened this issue Sep 9, 2022 · 5 comments · Fixed by #87817
Assignees
Labels
branch-master Failures and bugs on the master branch. C-test-failure Broken test (automatically or manually discovered). GA-blocker O-roachtest O-robot Originated from a bot. release-blocker Indicates a release-blocker. Use with branch-release-2x.x label to denote which branch is blocked. T-cdc
Milestone

Comments

@cockroach-teamcity
Copy link
Member

cockroach-teamcity commented Sep 9, 2022

roachtest.cdc/cloud-sink-gcs/rangefeed=true failed with artifacts on master @ a82711442c65cf14489c55041b45b11a1e38415b:

		  | _elapsed___errors__ops/sec(inst)___ops/sec(cum)__p50(ms)__p95(ms)__p99(ms)_pMax(ms)
		  |   377.0s        0            2.0            1.0     50.3     56.6     56.6     56.6 delivery
		  |   377.0s        0           12.0           10.5     22.0     27.3     54.5     54.5 newOrder
		  |   377.0s        0            4.0            1.1      5.2      5.8      5.8      5.8 orderStatus
		  |   377.0s        0           16.0           10.7     12.6     17.8     19.9     19.9 payment
		  |   377.0s        0            1.0            1.0      9.4      9.4      9.4      9.4 stockLevel
		  |   378.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 delivery
		  |   378.0s        0            2.0           10.5     23.1     24.1     24.1     24.1 newOrder
		  |   378.0s        0            0.0            1.1      0.0      0.0      0.0      0.0 orderStatus
		  |   378.0s        0           13.0           10.7     13.1     16.8     19.9     19.9 payment
		  |   378.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 stockLevel
		Wraps: (4) COMMAND_PROBLEM
		Wraps: (5) Node 4. Command with error:
		  | ``````
		  | ./workload run tpcc --warehouses=50 --duration=30m  {pgurl:1-3}
		  | ``````
		Wraps: (6) exit status 1
		Error types: (1) *withstack.withStack (2) *errutil.withPrefix (3) *cluster.WithCommandDetails (4) errors.Cmd (5) *hintdetail.withDetail (6) *exec.ExitError

	monitor.go:127,cdc.go:300,cdc.go:773,test_runner.go:906: monitor failure: monitor task failed: t.Fatal() was called
		(1) attached stack trace
		  -- stack trace:
		  | main.(*monitorImpl).WaitE
		  | 	main/pkg/cmd/roachtest/monitor.go:115
		  | main.(*monitorImpl).Wait
		  | 	main/pkg/cmd/roachtest/monitor.go:123
		  | github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests.cdcBasicTest
		  | 	github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests/cdc.go:300
		  | github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests.registerCDC.func7
		  | 	github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests/cdc.go:773
		  | main.(*testRunner).runTest.func2
		  | 	main/pkg/cmd/roachtest/test_runner.go:906
		Wraps: (2) monitor failure
		Wraps: (3) attached stack trace
		  -- stack trace:
		  | main.(*monitorImpl).wait.func2
		  | 	main/pkg/cmd/roachtest/monitor.go:171
		Wraps: (4) monitor task failed
		Wraps: (5) attached stack trace
		  -- stack trace:
		  | main.init
		  | 	main/pkg/cmd/roachtest/monitor.go:80
		  | runtime.doInit
		  | 	GOROOT/src/runtime/proc.go:6340
		  | runtime.main
		  | 	GOROOT/src/runtime/proc.go:233
		  | runtime.goexit
		  | 	GOROOT/src/runtime/asm_amd64.s:1594
		Wraps: (6) t.Fatal() was called
		Error types: (1) *withstack.withStack (2) *errutil.withPrefix (3) *withstack.withStack (4) *errutil.withPrefix (5) *withstack.withStack (6) *errutil.leafError

Parameters: ROACHTEST_cloud=gce , ROACHTEST_cpu=16 , ROACHTEST_ssd=0

Help

See: roachtest README

See: How To Investigate (internal)

/cc @cockroachdb/cdc

This test on roachdash | Improve this report!

Jira issue: CRDB-19470

Epic CRDB-11732

@cockroach-teamcity cockroach-teamcity added branch-master Failures and bugs on the master branch. C-test-failure Broken test (automatically or manually discovered). O-roachtest O-robot Originated from a bot. release-blocker Indicates a release-blocker. Use with branch-release-2x.x label to denote which branch is blocked. labels Sep 9, 2022
@cockroach-teamcity cockroach-teamcity added this to the 22.2 milestone Sep 9, 2022
@blathers-crl blathers-crl bot added the T-cdc label Sep 9, 2022
@cockroach-teamcity
Copy link
Member Author

roachtest.cdc/cloud-sink-gcs/rangefeed=true failed with artifacts on master @ 389661e823c19f318fa07ec2278336262531692d:

		  |  1421.0s        0            0.0           10.4      0.0      0.0      0.0      0.0 newOrder
		  |  1421.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 orderStatus
		  |  1421.0s        0            0.0           10.4      0.0      0.0      0.0      0.0 payment
		  |  1421.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 stockLevel
		  |  1422.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 delivery
		  |  1422.0s        0            0.0           10.4      0.0      0.0      0.0      0.0 newOrder
		  |  1422.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 orderStatus
		  |  1422.0s        0            0.0           10.4      0.0      0.0      0.0      0.0 payment
		  |  1422.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 stockLevel
		  |  1423.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 delivery
		  |  1423.0s        0            0.0           10.4      0.0      0.0      0.0      0.0 newOrder
		  |  1423.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 orderStatus
		  |  1423.0s        0            0.0           10.4      0.0      0.0      0.0      0.0 payment
		  |  1423.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 stockLevel
		Wraps: (4) secondary error attachment
		  | UNCLASSIFIED_PROBLEM: context canceled
		  | (1) UNCLASSIFIED_PROBLEM
		  | Wraps: (2) Node 4. Command with error:
		  |   | ``````
		  |   | ./workload run tpcc --warehouses=50 --duration=30m  {pgurl:1-3}
		  |   | ``````
		  | Wraps: (3) context canceled
		  | Error types: (1) errors.Unclassified (2) *hintdetail.withDetail (3) *errors.errorString
		Wraps: (5) context canceled
		Error types: (1) *withstack.withStack (2) *errutil.withPrefix (3) *cluster.WithCommandDetails (4) *secondary.withSecondaryError (5) *errors.errorString

	monitor.go:127,cdc.go:300,cdc.go:773,test_runner.go:917: monitor failure: monitor task failed: read tcp 172.17.0.3:56656 -> 35.231.210.0:26257: read: connection reset by peer
		(1) attached stack trace
		  -- stack trace:
		  | main.(*monitorImpl).WaitE
		  | 	main/pkg/cmd/roachtest/monitor.go:115
		  | main.(*monitorImpl).Wait
		  | 	main/pkg/cmd/roachtest/monitor.go:123
		  | github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests.cdcBasicTest
		  | 	github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests/cdc.go:300
		  | github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests.registerCDC.func7
		  | 	github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests/cdc.go:773
		  | [...repeated from below...]
		Wraps: (2) monitor failure
		Wraps: (3) attached stack trace
		  -- stack trace:
		  | main.(*monitorImpl).wait.func2
		  | 	main/pkg/cmd/roachtest/monitor.go:171
		  | runtime.goexit
		  | 	GOROOT/src/runtime/asm_amd64.s:1594
		Wraps: (4) monitor task failed
		Wraps: (5) read tcp 172.17.0.3:56656 -> 35.231.210.0:26257
		Wraps: (6) read
		Wraps: (7) connection reset by peer
		Error types: (1) *withstack.withStack (2) *errutil.withPrefix (3) *withstack.withStack (4) *errutil.withPrefix (5) *net.OpError (6) *os.SyscallError (7) syscall.Errno

Parameters: ROACHTEST_cloud=gce , ROACHTEST_cpu=16 , ROACHTEST_ssd=0

Help

See: roachtest README

See: How To Investigate (internal)

This test on roachdash | Improve this report!

@cockroach-teamcity
Copy link
Member Author

roachtest.cdc/cloud-sink-gcs/rangefeed=true failed with artifacts on master @ bc2e47da0523b347c28cf024707e80cd35d6c98a:

		  |    52.0s        0            0.0            1.1      0.0      0.0      0.0      0.0 orderStatus
		  |    52.0s        0            1.0           11.9     13.1     13.1     13.1     13.1 payment
		  |    52.0s        0            0.0            1.2      0.0      0.0      0.0      0.0 stockLevel
		  | _elapsed___errors__ops/sec(inst)___ops/sec(cum)__p50(ms)__p95(ms)__p99(ms)_pMax(ms)
		  |    53.0s        0            0.0            1.4      0.0      0.0      0.0      0.0 delivery
		  |    53.0s        0            0.0            9.7      0.0      0.0      0.0      0.0 newOrder
		  |    53.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 orderStatus
		  |    53.0s        0            0.0           11.7      0.0      0.0      0.0      0.0 payment
		  |    53.0s        0            0.0            1.2      0.0      0.0      0.0      0.0 stockLevel
		  |    54.0s        0            0.0            1.4      0.0      0.0      0.0      0.0 delivery
		  |    54.0s        0            0.0            9.5      0.0      0.0      0.0      0.0 newOrder
		  |    54.0s        0            0.0            1.0      0.0      0.0      0.0      0.0 orderStatus
		  |    54.0s        0            1.0           11.5     11.5     11.5     11.5     11.5 payment
		  |    54.0s        0            0.0            1.2      0.0      0.0      0.0      0.0 stockLevel
		Wraps: (4) secondary error attachment
		  | UNCLASSIFIED_PROBLEM: context canceled
		  | (1) UNCLASSIFIED_PROBLEM
		  | Wraps: (2) Node 4. Command with error:
		  |   | ``````
		  |   | ./workload run tpcc --warehouses=50 --duration=30m  {pgurl:1-3}
		  |   | ``````
		  | Wraps: (3) context canceled
		  | Error types: (1) errors.Unclassified (2) *hintdetail.withDetail (3) *errors.errorString
		Wraps: (5) context canceled
		Error types: (1) *withstack.withStack (2) *errutil.withPrefix (3) *cluster.WithCommandDetails (4) *secondary.withSecondaryError (5) *errors.errorString

	monitor.go:127,cdc.go:300,cdc.go:773,test_runner.go:917: monitor failure: monitor task failed: read tcp 172.17.0.3:44836 -> 34.139.109.27:26257: read: connection reset by peer
		(1) attached stack trace
		  -- stack trace:
		  | main.(*monitorImpl).WaitE
		  | 	main/pkg/cmd/roachtest/monitor.go:115
		  | main.(*monitorImpl).Wait
		  | 	main/pkg/cmd/roachtest/monitor.go:123
		  | github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests.cdcBasicTest
		  | 	github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests/cdc.go:300
		  | github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests.registerCDC.func7
		  | 	github.com/cockroachdb/cockroach/pkg/cmd/roachtest/tests/cdc.go:773
		  | [...repeated from below...]
		Wraps: (2) monitor failure
		Wraps: (3) attached stack trace
		  -- stack trace:
		  | main.(*monitorImpl).wait.func2
		  | 	main/pkg/cmd/roachtest/monitor.go:171
		  | runtime.goexit
		  | 	GOROOT/src/runtime/asm_amd64.s:1594
		Wraps: (4) monitor task failed
		Wraps: (5) read tcp 172.17.0.3:44836 -> 34.139.109.27:26257
		Wraps: (6) read
		Wraps: (7) connection reset by peer
		Error types: (1) *withstack.withStack (2) *errutil.withPrefix (3) *withstack.withStack (4) *errutil.withPrefix (5) *net.OpError (6) *os.SyscallError (7) syscall.Errno

Parameters: ROACHTEST_cloud=gce , ROACHTEST_cpu=16 , ROACHTEST_ssd=0

Help

See: roachtest README

See: How To Investigate (internal)

This test on roachdash | Improve this report!

@cockroach-teamcity
Copy link
Member Author

roachtest.cdc/cloud-sink-gcs/rangefeed=true failed with artifacts on master @ 773568fbda06ba9be9fb1bc34a331f21c8891ffa:

test artifacts and logs in: /artifacts/cdc/cloud-sink-gcs/rangefeed=true/run_1
	cdc.go:1734,cdc.go:302,cdc.go:773,test_runner.go:917: initial scan did not complete

Parameters: ROACHTEST_cloud=gce , ROACHTEST_cpu=16 , ROACHTEST_ssd=0

Help

See: roachtest README

See: How To Investigate (internal)

This test on roachdash | Improve this report!

@miretskiy
Copy link
Contributor

All but the first one have the following:

W220911 06:28:02.754456 8411 ccl/changefeedccl/changefeed_stmt.go:1014 ⋮ [n1,job=795630898610962433] 873  WARNING: CHANGEFEED job 795630898610962433 encountered retryable error: ‹retryable changefeed error›: ‹upload request to https://storage.googleapis.com/upload/storage/v1/b/cockroac
h-tmp/o?alt=json&name=roachtest%2F20220911062743%2F2022-09-11%2F202209110628010617934210000000000-bcf0778e4a8893aa-2-2-00000000-customer-b.ndjson&prettyPrint=false&projection=full&uploadType=resumable&upload_id=ADPycdshKvxSgiWngFkybzT2dWJstIr1ZF-QoQn5YwXJs-vCjfdwKAUwZ5dNsF373zPDEIsX0IybtW9MbLXsALnSpAAUXehqBuRH not sent, choose larger value for ChunkRetryDealine›

ac2387e might be suspect
(though not clear why that is).

@adityamaru investigating.

@miretskiy miretskiy added GA-blocker release-blocker Indicates a release-blocker. Use with branch-release-2x.x label to denote which branch is blocked. and removed release-blocker Indicates a release-blocker. Use with branch-release-2x.x label to denote which branch is blocked. labels Sep 12, 2022
craig bot pushed a commit that referenced this issue Sep 12, 2022
87817: gcp: fix ChunkRetryDeadline default value r=miretskiy a=adityamaru

A previous patch made the GCP ChunkRetryDeadline configurable but incorrectly set the default to a very small value instead of 60seconds.

Fixes: #87677

Release note (bug fix): fix incorrect default value of `cloudstorage.gs.chunking.retry_timeout` to 60 seconds

Co-authored-by: adityamaru <[email protected]>
@adityamaru
Copy link
Contributor

ac2387e was indeed the issue. The default value of the chunk retry deadline was set to a very small value of 60ms, which is why every upload to GCP was being rejected. This #87817 should fix it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Labels
branch-master Failures and bugs on the master branch. C-test-failure Broken test (automatically or manually discovered). GA-blocker O-roachtest O-robot Originated from a bot. release-blocker Indicates a release-blocker. Use with branch-release-2x.x label to denote which branch is blocked. T-cdc
Projects
None yet
Development

Successfully merging a pull request may close this issue.

3 participants