From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8FBDA413782; Sat, 3 Oct 2026 10:54:37 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791024949; cv=none; b=hMFqWagwJqnhrhfmo5fy1q+LBHJVjAZq5Bqeak2lxCCeBCOihEfyg18kQ6gcDP0niJfsmgFNKsBD2DAkNgScmnMc+H6KtArpFiD+sm24xRxvQH6wpzza2yPQrGRyrhjyyAMwlWrchqKu4OMd6XUZp2O7DwoVY5E/oE/mATR/Emc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791024949; c=relaxed/simple; bh=lN8g7l4Jho3tRHMchMhE/PY/TaGuHFfx0eGb78Z+g/g=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=FbUYoeZWKb7dU2LobQ2UYhzkCTJVTSn2RIVVSuuNMmcLmIZ2/h4XwVdPgPaow04QN9P4dqedFawINt6QTb+Z7ZtUVC0OJvry584VYCddJTGfTD9zNnBWzuuIXJiYEPM6UJ+Drkfo6KtClUtSUWKafSwiXof+2FfbwAUF3rzQD28= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b=wbiZ/AOa; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linuxfoundation.org header.i=@linuxfoundation.org header.b="wbiZ/AOa" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 4BAFD1F0089D; Sat, 3 Oct 2026 10:54:33 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linuxfoundation.org; s=korg; t=1791024874; bh=Jr43APnITqP0R0J6iqGBaDQJS1ZXRAhZ+ebokbTcQyM=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=wbiZ/AOayEUndYKvbLAanYBBCfKa125nSIs9ETn6IRGJOxua10LIAbKphEbbhwk+y esqd0CtVwtYZu23GOGaFYzPn28sRMf24mSTP2HjjoViZuPWezKhC6Gv9aGqZBVUxFm ji3l4BBwaDRuE0g72PErJOgnzs+XMCVydaCmlm5E= From: Greg Kroah-Hartman To: linux-kernel@vger.kernel.org, akpm@linux-foundation.org, torvalds@linux-foundation.org, stable@vger.kernel.org Cc: lwn@lwn.net, jslaby@suse.cz, Greg Kroah-Hartman Subject: Re: Linux 7.2.9 Date: Sat, 3 Oct 2026 12:54:23 +0200 Message-ID: <2026100323-snarl-appetite-315f@gregkh> X-Mailer: git-send-email 2.56.0 In-Reply-To: <2026100323-uncombed-july-abe2@gregkh> References: <2026100323-uncombed-july-abe2@gregkh> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit diff --git a/Documentation/arch/arm64/booting.rst b/Documentation/arch/arm64/booting.rst index 13ef311dace8..3fea4b14ef7c 100644 --- a/Documentation/arch/arm64/booting.rst +++ b/Documentation/arch/arm64/booting.rst @@ -465,6 +465,7 @@ Before jumping into the kernel, the following conditions must be met: - HDFGWTR2_EL2.nPMICNTR_EL0 (bit 2) must be initialised to 0b1. - HDFGWTR2_EL2.nPMICFILTR_EL0 (bit 3) must be initialised to 0b1. - HDFGWTR2_EL2.nPMUACR_EL1 (bit 4) must be initialised to 0b1. + - HDFGWTR2_EL2.nPMZR_EL0 (bit 21) must be initialised to 0b1. For CPUs with SPE data source filtering (FEAT_SPE_FDS): diff --git a/Documentation/arch/s390/pci.rst b/Documentation/arch/s390/pci.rst index 80f4ba193159..565434626fb5 100644 --- a/Documentation/arch/s390/pci.rst +++ b/Documentation/arch/s390/pci.rst @@ -67,7 +67,7 @@ Entries specific to zPCI functions and entries that hold zPCI information. A physical function that currently supports a virtual function cannot be powered off until all virtual functions are removed with: - echo 0 > /sys/bus/pci/devices/DDDD:BB:dd.f/sriov_numvf + echo 0 > /sys/bus/pci/devices/DDDD:BB:dd.f/sriov_numvfs * /sys/bus/pci/devices/DDDD:BB:dd.f/: diff --git a/Documentation/devicetree/bindings/pinctrl/qcom,nord-tlmm.yaml b/Documentation/devicetree/bindings/pinctrl/qcom,nord-tlmm.yaml index 4bb511719f31..56758a85ead8 100644 --- a/Documentation/devicetree/bindings/pinctrl/qcom,nord-tlmm.yaml +++ b/Documentation/devicetree/bindings/pinctrl/qcom,nord-tlmm.yaml @@ -67,7 +67,7 @@ $defs: Specify the alternative function to be configured for the specified pins. - enum: [ aoss_cti, atest_char, atest_usb20, atest_usb21, + enum: [ gpio, aoss_cti, atest_char, atest_usb20, atest_usb21, aud_intfc0_clk, aud_intfc0_data, aud_intfc0_ws, aud_intfc10_clk, aud_intfc10_data, aud_intfc10_ws, aud_intfc1_clk, aud_intfc1_data, aud_intfc1_ws, @@ -98,9 +98,10 @@ $defs: pcie3_clk_req_n, phase_flag, pll_bist_sync, pll_clk_aux, prng_rosc0, prng_rosc1, pwrbrk_i_n, qdss, qdss_cti, qspi, qup0_se0, qup0_se1, qup0_se2, qup0_se3, qup0_se4, qup0_se5, - qup1_se0, qup1_se1, qup1_se3, qup1_se2, qup1_se4, qup1_se5, + qup1_se0, qup1_se1, qup1_se2_01, qup1_se2_23, qup1_se3_01, + qup1_se3_23, qup1_se4, qup1_se5, qup1_se6, qup2_se0, qup2_se1, qup2_se2, qup2_se3, qup2_se4, - qup2_se5, qup2_se6, + qup2_se5, qup2_se6, qup3_se0_mira, qup3_se0_mirb, sailss_ospi, sdc4_clk, sdc4_cmd, sdc4_data, smb_alert, smb_alert_n, smb_clk, smb_dat, tb_trig_sdc4, tmess_prng0, tmess_prng1, tsc_timer, tsense_pwm, usb0_hs, diff --git a/Makefile b/Makefile index ad1a1ffe789e..6f8ed7329a20 100644 --- a/Makefile +++ b/Makefile @@ -1,7 +1,7 @@ # SPDX-License-Identifier: GPL-2.0 VERSION = 7 PATCHLEVEL = 2 -SUBLEVEL = 8 +SUBLEVEL = 9 EXTRAVERSION = NAME = Baby Opossum Posse diff --git a/arch/arm64/include/asm/el2_setup.h b/arch/arm64/include/asm/el2_setup.h index aa8ec9df8024..87560d8b254e 100644 --- a/arch/arm64/include/asm/el2_setup.h +++ b/arch/arm64/include/asm/el2_setup.h @@ -418,6 +418,7 @@ b.lt .Lskip_fgt2_\@ mov x0, xzr + mov x2, xzr mrs x1, id_aa64dfr0_el1 ubfx x1, x1, #ID_AA64DFR0_EL1_PMUVer_SHIFT, #4 cmp x1, #ID_AA64DFR0_EL1_PMUVer_V3P9 @@ -426,6 +427,11 @@ orr x0, x0, #HDFGRTR2_EL2_nPMICNTR_EL0 orr x0, x0, #HDFGRTR2_EL2_nPMICFILTR_EL0 orr x0, x0, #HDFGRTR2_EL2_nPMUACR_EL1 + orr x2, x2, #HDFGWTR2_EL2_nPMICNTR_EL0 + orr x2, x2, #HDFGWTR2_EL2_nPMICFILTR_EL0 + orr x2, x2, #HDFGWTR2_EL2_nPMUACR_EL1 + /* PMZR_EL0 is write-only, so it has no read trap to disable */ + orr x2, x2, #HDFGWTR2_EL2_nPMZR_EL0 .Lskip_pmuv3p9_\@: /* If SPE is implemented, */ __spe_vers_imp .Lskip_spefds_\@, ID_AA64DFR0_EL1_PMSVer_IMP, x1 @@ -436,10 +442,11 @@ cbz x1, .Lskip_spefds_\@ /* disable traps of PMSDSFR to EL2. */ orr x0, x0, #HDFGRTR2_EL2_nPMSDSFR_EL1 + orr x2, x2, #HDFGWTR2_EL2_nPMSDSFR_EL1 .Lskip_spefds_\@: msr_s SYS_HDFGRTR2_EL2, x0 - msr_s SYS_HDFGWTR2_EL2, x0 + msr_s SYS_HDFGWTR2_EL2, x2 msr_s SYS_HFGRTR2_EL2, xzr msr_s SYS_HFGWTR2_EL2, xzr msr_s SYS_HFGITR2_EL2, xzr diff --git a/arch/arm64/include/asm/io.h b/arch/arm64/include/asm/io.h index 21c8e400107c..31e67f6bd4a7 100644 --- a/arch/arm64/include/asm/io.h +++ b/arch/arm64/include/asm/io.h @@ -272,7 +272,8 @@ static inline void __iomem *ioremap_prot(phys_addr_t phys, size_t size, pgprot_t prot; ptval_t user_prot_val = pgprot_val(user_prot); - if (WARN_ON_ONCE(!(user_prot_val & PTE_USER))) + /* Reject PROT_NONE and exec-only */ + if (!(user_prot_val & PTE_USER)) return NULL; prot = __pgprot_modify(PAGE_KERNEL, PTE_ATTRINDX_MASK, diff --git a/arch/arm64/include/asm/kvm_host.h b/arch/arm64/include/asm/kvm_host.h index ac16f96c878d..7f94635e69c9 100644 --- a/arch/arm64/include/asm/kvm_host.h +++ b/arch/arm64/include/asm/kvm_host.h @@ -322,7 +322,7 @@ struct kvm_arch { * Stage 2 paging state for VMs with nested S2 using a virtual * VMID. */ - struct kvm_s2_mmu *nested_mmus; + struct kvm_s2_mmu **nested_mmus; size_t nested_mmus_size; int nested_mmus_next; diff --git a/arch/arm64/include/asm/kvm_nested.h b/arch/arm64/include/asm/kvm_nested.h index 1ed708335809..586026e85903 100644 --- a/arch/arm64/include/asm/kvm_nested.h +++ b/arch/arm64/include/asm/kvm_nested.h @@ -66,7 +66,8 @@ static inline u64 translate_ttbr0_el2_to_ttbr0_el1(u64 ttbr0) extern bool forward_smc_trap(struct kvm_vcpu *vcpu); extern bool forward_debug_exception(struct kvm_vcpu *vcpu); -extern void kvm_init_nested(struct kvm *kvm); +extern int kvm_init_nested(struct kvm *kvm); +extern void kvm_destroy_nested(struct kvm *kvm); extern int kvm_vcpu_init_nested(struct kvm_vcpu *vcpu); extern void kvm_init_nested_s2_mmu(struct kvm_s2_mmu *mmu); extern struct kvm_s2_mmu *lookup_s2_mmu(struct kvm_vcpu *vcpu); diff --git a/arch/arm64/include/asm/kvm_pkvm.h b/arch/arm64/include/asm/kvm_pkvm.h index 74fedd9c5ff0..ce5fd7cfc4e4 100644 --- a/arch/arm64/include/asm/kvm_pkvm.h +++ b/arch/arm64/include/asm/kvm_pkvm.h @@ -62,8 +62,7 @@ static inline bool kvm_pkvm_ioctl_allowed(struct kvm *kvm, unsigned int ioctl) int r; r = kvm_get_cap_for_kvm_ioctl(ioctl, &ext); - - if (WARN_ON_ONCE(r < 0)) + if (r < 0) return false; return kvm_pkvm_ext_allowed(kvm, ext); diff --git a/arch/arm64/kernel/cpu_errata.c b/arch/arm64/kernel/cpu_errata.c index b0937e3f6fb4..be47908bb9d2 100644 --- a/arch/arm64/kernel/cpu_errata.c +++ b/arch/arm64/kernel/cpu_errata.c @@ -28,18 +28,21 @@ bool cpu_errata_set_target_impl(u64 num, void *impl_cpus) return true; } +static inline bool __is_midr_in_range(u32 midr, struct midr_range const *range) +{ + return midr_is_cpu_model_range(midr, range->model, + range->rv_min, range->rv_max); +} + static inline bool is_midr_in_range(struct midr_range const *range) { int i; if (!target_impl_cpu_num) - return midr_is_cpu_model_range(read_cpuid_id(), range->model, - range->rv_min, range->rv_max); + return __is_midr_in_range(read_cpuid_id(), range); for (i = 0; i < target_impl_cpu_num; i++) { - if (midr_is_cpu_model_range(target_impl_cpus[i].midr, - range->model, - range->rv_min, range->rv_max)) + if (__is_midr_in_range(target_impl_cpus[i].midr, range)) return true; } return false; @@ -59,7 +62,7 @@ __is_affected_midr_range(const struct arm64_cpu_capabilities *entry, u32 midr, u32 revidr) { const struct arm64_midr_revidr *fix; - if (!is_midr_in_range(&entry->midr_range)) + if (!__is_midr_in_range(midr, &entry->midr_range)) return false; midr &= MIDR_REVISION_MASK | MIDR_VARIANT_MASK; diff --git a/arch/arm64/kvm/arm.c b/arch/arm64/kvm/arm.c index 9a6c72a18672..ad04ef4bb821 100644 --- a/arch/arm64/kvm/arm.c +++ b/arch/arm64/kvm/arm.c @@ -236,8 +236,6 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type) mutex_unlock(&kvm->lock); #endif - kvm_init_nested(kvm); - ret = kvm_share_hyp(kvm, kvm + 1); if (ret) return ret; @@ -252,6 +250,10 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type) if (ret) goto err_free_cpumask; + ret = kvm_init_nested(kvm); + if (ret) + goto err_uninit_mmu; + if (is_protected_kvm_enabled()) { /* * If any failures occur after this is successful, make sure to @@ -280,6 +282,7 @@ int kvm_arch_init_vm(struct kvm *kvm, unsigned long type) err_uninit_mmu: kvm_uninit_stage2_mmu(kvm); + kvm_destroy_nested(kvm); err_free_cpumask: free_cpumask_var(kvm->arch.supported_cpus); err_unshare_kvm: @@ -337,6 +340,7 @@ void kvm_arch_destroy_vm(struct kvm *kvm) kvm_unshare_hyp(kvm, kvm + 1); + kvm_destroy_nested(kvm); kvm_arm_teardown_hypercalls(kvm); } diff --git a/arch/arm64/kvm/emulate-nested.c b/arch/arm64/kvm/emulate-nested.c index 3c82f392845d..b32742d9dd73 100644 --- a/arch/arm64/kvm/emulate-nested.c +++ b/arch/arm64/kvm/emulate-nested.c @@ -1432,7 +1432,7 @@ static const struct encoding_to_trap_config encoding_to_fgt[] __initconst = { SR_FGT(OP_AT_S1E1A, HFGITR, ATS1E1A, 1), SR_FGT(OP_COSP_RCTX, HFGITR, COSPRCTX, 1), SR_FGT(OP_GCSPUSHX, HFGITR, nGCSEPP, 0), - SR_FGT(OP_GCSPOPX, HFGITR, nGCSEPP, 0), + SR_FGT(OP_GCSPOPCX, HFGITR, nGCSEPP, 0), SR_FGT(OP_GCSPUSHM, HFGITR, nGCSPUSHM_EL1, 0), SR_FGT(OP_BRB_IALL, HFGITR, nBRBIALL, 0), SR_FGT(OP_BRB_INJ, HFGITR, nBRBINJ, 0), diff --git a/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h b/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h index 29935c7da1de..cab27f7bd423 100644 --- a/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h +++ b/arch/arm64/kvm/hyp/include/nvhe/mem_protect.h @@ -52,6 +52,7 @@ int __pkvm_host_test_clear_young_guest(u64 gfn, u64 nr_pages, bool mkold, struct int __pkvm_host_mkyoung_guest(u64 gfn, struct pkvm_hyp_vcpu *vcpu); bool addr_is_memory(phys_addr_t phys); +bool addr_is_hyp_text(phys_addr_t phys); int host_stage2_idmap_locked(phys_addr_t addr, u64 size, enum kvm_pgtable_prot prot); int host_stage2_set_owner_locked(phys_addr_t addr, u64 size, u8 owner_id); int kvm_host_prepare_stage2(void *pgt_pool_base); diff --git a/arch/arm64/kvm/hyp/nvhe/mem_protect.c b/arch/arm64/kvm/hyp/nvhe/mem_protect.c index 4e329e39a695..7a9936130642 100644 --- a/arch/arm64/kvm/hyp/nvhe/mem_protect.c +++ b/arch/arm64/kvm/hyp/nvhe/mem_protect.c @@ -443,6 +443,14 @@ bool addr_is_memory(phys_addr_t phys) return !!find_mem_range(phys, &range); } +bool addr_is_hyp_text(phys_addr_t phys) +{ + phys_addr_t start = ALIGN_DOWN(__hyp_pa(__hyp_text_start), PAGE_SIZE); + phys_addr_t end = PAGE_ALIGN(__hyp_pa(__hyp_text_end)); + + return phys >= start && phys < end; +} + static bool is_in_mem_range(u64 addr, struct kvm_mem_range *range) { return range->start <= addr && addr < range->end; diff --git a/arch/arm64/kvm/hyp/nvhe/pkvm.c b/arch/arm64/kvm/hyp/nvhe/pkvm.c index 24d6f164129a..53532f0c4162 100644 --- a/arch/arm64/kvm/hyp/nvhe/pkvm.c +++ b/arch/arm64/kvm/hyp/nvhe/pkvm.c @@ -360,7 +360,7 @@ static void pkvm_init_features_from_host(struct pkvm_hyp_vm *hyp_vm, const struc if (test_bit(KVM_ARCH_FLAG_WRITABLE_IMP_ID_REGS, &host_arch_flags)) hyp_vm->kvm.arch.midr_el1 = host_kvm->arch.midr_el1; - return; + goto out; } if (kvm_pkvm_ext_allowed(kvm, KVM_CAP_ARM_MTE)) @@ -379,13 +379,14 @@ static void pkvm_init_features_from_host(struct pkvm_hyp_vm *hyp_vm, const struc if (kvm_pkvm_ext_allowed(kvm, KVM_CAP_ARM_PTRAUTH_GENERIC)) set_bit(KVM_ARM_VCPU_PTRAUTH_GENERIC, allowed_features); - if (kvm_pkvm_ext_allowed(kvm, KVM_CAP_ARM_SVE)) { + if (kvm_pkvm_ext_allowed(kvm, KVM_CAP_ARM_SVE)) set_bit(KVM_ARM_VCPU_SVE, allowed_features); - kvm->arch.flags |= host_arch_flags & BIT(KVM_ARCH_FLAG_GUEST_HAS_SVE); - } bitmap_and(kvm->arch.vcpu_features, host_kvm->arch.vcpu_features, allowed_features, KVM_VCPU_MAX_FEATURES); +out: + __assign_bit(KVM_ARCH_FLAG_GUEST_HAS_SVE, &kvm->arch.flags, + kvm_vcpu_has_feature(kvm, KVM_ARM_VCPU_SVE)); } static void unpin_host_vcpu(struct kvm_vcpu *host_vcpu) @@ -451,7 +452,7 @@ static int pkvm_vcpu_init_sve(struct pkvm_hyp_vcpu *hyp_vcpu, struct kvm_vcpu *h unsigned int sve_max_vl; size_t sve_state_size; void *sve_state; - int ret = 0; + int ret; if (!vcpu_has_feature(vcpu, KVM_ARM_VCPU_SVE)) { vcpu_clear_flag(vcpu, VCPU_SVE_FINALIZED); @@ -460,25 +461,21 @@ static int pkvm_vcpu_init_sve(struct pkvm_hyp_vcpu *hyp_vcpu, struct kvm_vcpu *h /* Limit guest vector length to the maximum supported by the host. */ sve_max_vl = min(READ_ONCE(host_vcpu->arch.sve_max_vl), kvm_host_sve_max_vl); - sve_state_size = sve_state_size_from_vl(sve_max_vl); sve_state = kern_hyp_va(READ_ONCE(host_vcpu->arch.sve_state)); - if (!sve_state || !sve_state_size) { - ret = -EINVAL; - goto err; - } + if (!sve_vl_valid(sve_max_vl) || !sve_state) + return -EINVAL; + + sve_state_size = sve_state_size_from_vl(sve_max_vl); ret = hyp_pin_shared_mem(sve_state, sve_state + sve_state_size); if (ret) - goto err; + return ret; vcpu->arch.sve_state = sve_state; vcpu->arch.sve_max_vl = sve_max_vl; return 0; -err: - clear_bit(KVM_ARM_VCPU_SVE, vcpu->kvm->arch.vcpu_features); - return ret; } static int vm_copy_id_regs(struct pkvm_hyp_vcpu *hyp_vcpu) diff --git a/arch/arm64/kvm/hyp/nvhe/setup.c b/arch/arm64/kvm/hyp/nvhe/setup.c index 75b00c323310..bb667cd7080b 100644 --- a/arch/arm64/kvm/hyp/nvhe/setup.c +++ b/arch/arm64/kvm/hyp/nvhe/setup.c @@ -217,7 +217,7 @@ static int fix_host_ownership_walker(const struct kvm_pgtable_visit_ctx *ctx, case PKVM_PAGE_OWNED: set_hyp_state(page, PKVM_PAGE_OWNED); /* hyp text is RO in the host stage-2 to be inspected on panic. */ - if (prot == PAGE_HYP_EXEC) { + if (addr_is_hyp_text(phys)) { set_host_state(page, PKVM_NOPAGE); return host_stage2_idmap_locked(phys, PAGE_SIZE, KVM_PGTABLE_PROT_R); } else { @@ -269,6 +269,16 @@ static int fix_host_ownership(void) return ret; } + /* The stacks sit in the private VA range, not the linear map. */ + for (i = 0; i < hyp_nr_cpus; i++) { + struct kvm_nvhe_init_params *params = per_cpu_ptr(&kvm_init_params, i); + u64 start = params->stack_hyp_va - NVHE_STACK_SIZE; + + ret = kvm_pgtable_walk(&pkvm_pgtable, start, NVHE_STACK_SIZE, &walker); + if (ret) + return ret; + } + return 0; } diff --git a/arch/arm64/kvm/hypercalls.c b/arch/arm64/kvm/hypercalls.c index b11b8821c9fb..dfa25bb6f25d 100644 --- a/arch/arm64/kvm/hypercalls.c +++ b/arch/arm64/kvm/hypercalls.c @@ -185,7 +185,8 @@ static int kvm_smccc_set_filter(struct kvm *kvm, struct kvm_smccc_filter __user start = filter.base; end = start + filter.nr_functions - 1; - if (end < start || filter.action >= NR_SMCCC_FILTER_ACTIONS) + if (!filter.nr_functions || end < start || + filter.action >= NR_SMCCC_FILTER_ACTIONS) return -EINVAL; mutex_lock(&kvm->arch.config_lock); diff --git a/arch/arm64/kvm/mmu.c b/arch/arm64/kvm/mmu.c index 2d95203386ba..2ed0836cb929 100644 --- a/arch/arm64/kvm/mmu.c +++ b/arch/arm64/kvm/mmu.c @@ -59,27 +59,36 @@ static phys_addr_t stage2_range_addr_end(phys_addr_t addr, phys_addr_t end) * long will also starve other vCPUs. We have to also make sure that the page * tables are not freed while we released the lock. */ -static int stage2_apply_range(struct kvm_s2_mmu *mmu, phys_addr_t addr, +static int stage2_apply_range(struct kvm_s2_mmu *mmu, phys_addr_t start, phys_addr_t end, int (*fn)(struct kvm_pgtable *, u64, u64), bool resched) { struct kvm *kvm = kvm_s2_mmu_to_kvm(mmu); + bool lock_dropped = false; + phys_addr_t addr = start; int ret; u64 next; do { struct kvm_pgtable *pgt = mmu->pgt; + /* + * We may be raced on PGT teardown when we release the + * kvm->mmu_lock. That's fine as the PGT is legitimately no + * longer present. + */ if (!pgt) - return -EINVAL; + return lock_dropped ? 0 : -EINVAL; next = stage2_range_addr_end(addr, end); ret = fn(pgt, addr, next - addr); if (ret) break; - if (resched && next != end) + if (resched && next != end) { cond_resched_rwlock_write(&kvm->mmu_lock); + lock_dropped = true; + } } while (addr = next, addr != end); return ret; diff --git a/arch/arm64/kvm/nested.c b/arch/arm64/kvm/nested.c index 550c9bd3dbe7..982bbf3110e0 100644 --- a/arch/arm64/kvm/nested.c +++ b/arch/arm64/kvm/nested.c @@ -44,10 +44,23 @@ struct vncr_tlb { */ #define S2_MMU_PER_VCPU 2 -void kvm_init_nested(struct kvm *kvm) +int kvm_init_nested(struct kvm *kvm) { - kvm->arch.nested_mmus = NULL; + kvm->arch.nested_mmus = kvmalloc_objs(struct kvm_s2_mmu *, + KVM_MAX_VCPUS * S2_MMU_PER_VCPU, + GFP_KERNEL_ACCOUNT); kvm->arch.nested_mmus_size = 0; + + return kvm->arch.nested_mmus ? 0 : -ENOMEM; +} + +void kvm_destroy_nested(struct kvm *kvm) +{ + for (int i = 0; i < kvm->arch.nested_mmus_size; i+= S2_MMU_PER_VCPU) + kvfree(kvm->arch.nested_mmus[i]); + + kvm->arch.nested_mmus_size = 0; + kvfree(kvm->arch.nested_mmus); } static int init_nested_s2_mmu(struct kvm *kvm, struct kvm_s2_mmu *mmu) @@ -68,8 +81,9 @@ static int init_nested_s2_mmu(struct kvm *kvm, struct kvm_s2_mmu *mmu) int kvm_vcpu_init_nested(struct kvm_vcpu *vcpu) { struct kvm *kvm = vcpu->kvm; - struct kvm_s2_mmu *tmp; - int num_mmus, ret = 0; + int num_mmus; + + lockdep_assert_held(&kvm->arch.config_lock); if (test_bit(KVM_ARM_VCPU_HAS_EL2_E2H0, kvm->arch.vcpu_features) && !cpus_have_final_cap(ARM64_HAS_HCR_NV1)) @@ -82,51 +96,40 @@ int kvm_vcpu_init_nested(struct kvm_vcpu *vcpu) if (!vcpu->arch.ctxt.vncr_array) return -ENOMEM; - /* - * Let's treat memory allocation failures as benign: If we fail to - * allocate anything, return an error and keep the allocated array - * alive. Userspace may try to recover by initializing the vcpu - * again, and there is no reason to affect the whole VM for this. - */ num_mmus = atomic_read(&kvm->online_vcpus) * S2_MMU_PER_VCPU; if (num_mmus > kvm->arch.nested_mmus_size) { - tmp = kvcalloc(num_mmus, sizeof(*tmp), GFP_KERNEL_ACCOUNT); - if (!tmp) - return -ENOMEM; + struct kvm_s2_mmu *tmp; + int i, ret = 0; - write_lock(&kvm->mmu_lock); - - if (kvm->arch.nested_mmus_size) { - memcpy(tmp, kvm->arch.nested_mmus, - size_mul(sizeof(*tmp), kvm->arch.nested_mmus_size)); + tmp = kvcalloc(S2_MMU_PER_VCPU, sizeof(*tmp), GFP_KERNEL_ACCOUNT); + if (!tmp) + ret = -ENOMEM; - for (int i = 0; i < kvm->arch.nested_mmus_size; i++) - tmp[i].pgt->mmu = &tmp[i]; + for (i = 0; !ret && i < S2_MMU_PER_VCPU; i++) { + ret = init_nested_s2_mmu(kvm, &tmp[i]); + if (ret) + break; } - swap(kvm->arch.nested_mmus, tmp); - - write_unlock(&kvm->mmu_lock); - - kvfree(tmp); - } + if (ret) { + while (--i >= 0) + kvm_free_stage2_pgd(&tmp[i]); - for (int i = kvm->arch.nested_mmus_size; !ret && i < num_mmus; i++) - ret = init_nested_s2_mmu(kvm, &kvm->arch.nested_mmus[i]); + kvfree(tmp); + free_page((unsigned long)vcpu->arch.ctxt.vncr_array); + vcpu->arch.ctxt.vncr_array = NULL; + return ret; + } - if (ret) { - for (int i = kvm->arch.nested_mmus_size; i < num_mmus; i++) - kvm_free_stage2_pgd(&kvm->arch.nested_mmus[i]); + guard(write_lock)(&kvm->mmu_lock); - free_page((unsigned long)vcpu->arch.ctxt.vncr_array); - vcpu->arch.ctxt.vncr_array = NULL; + for (i = 0; i < S2_MMU_PER_VCPU; i++) + kvm->arch.nested_mmus[i + kvm->arch.nested_mmus_size] = &tmp[i]; - return ret; + kvm->arch.nested_mmus_size += S2_MMU_PER_VCPU; } - kvm->arch.nested_mmus_size = num_mmus; - return 0; } @@ -740,7 +743,7 @@ void kvm_s2_mmu_iterate_by_vmid(struct kvm *kvm, u16 vmid, write_lock(&kvm->mmu_lock); for (int i = 0; i < kvm->arch.nested_mmus_size; i++) { - struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i]; + struct kvm_s2_mmu *mmu = kvm->arch.nested_mmus[i]; if (!kvm_s2_mmu_valid(mmu)) continue; @@ -782,7 +785,7 @@ struct kvm_s2_mmu *lookup_s2_mmu(struct kvm_vcpu *vcpu) * if S2 translation is disabled. */ for (int i = 0; i < kvm->arch.nested_mmus_size; i++) { - struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i]; + struct kvm_s2_mmu *mmu = kvm->arch.nested_mmus[i]; if (!kvm_s2_mmu_valid(mmu)) continue; @@ -821,7 +824,7 @@ static struct kvm_s2_mmu *get_s2_mmu_nested(struct kvm_vcpu *vcpu) for (i = kvm->arch.nested_mmus_next; i < (kvm->arch.nested_mmus_size + kvm->arch.nested_mmus_next); i++) { - s2_mmu = &kvm->arch.nested_mmus[i % kvm->arch.nested_mmus_size]; + s2_mmu = kvm->arch.nested_mmus[i % kvm->arch.nested_mmus_size]; if (atomic_read(&s2_mmu->refcnt) == 0) break; @@ -1256,6 +1259,17 @@ void kvm_handle_s1e2_tlbi(struct kvm_vcpu *vcpu, u32 inst, u64 val) invalidate_vncr_va(vcpu->kvm, &scope); } +static void kvm_invalidate_vncr_ipa_all(struct kvm *kvm) +{ + struct kvm_pgtable *pgt = kvm->arch.mmu.pgt; + + lockdep_assert_held_write(&kvm->mmu_lock); + + /* if the mmu lock was dropped, pgt teardown may have raced. */ + if (pgt) + kvm_invalidate_vncr_ipa(kvm, 0, BIT(pgt->ia_bits)); +} + void kvm_nested_s2_wp(struct kvm *kvm) { int i; @@ -1266,13 +1280,13 @@ void kvm_nested_s2_wp(struct kvm *kvm) return; for (i = 0; i < kvm->arch.nested_mmus_size; i++) { - struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i]; + struct kvm_s2_mmu *mmu = kvm->arch.nested_mmus[i]; if (kvm_s2_mmu_valid(mmu)) kvm_stage2_wp_range(mmu, 0, kvm_phys_size(mmu)); } - kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits)); + kvm_invalidate_vncr_ipa_all(kvm); } void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block) @@ -1285,13 +1299,13 @@ void kvm_nested_s2_unmap(struct kvm *kvm, bool may_block) return; for (i = 0; i < kvm->arch.nested_mmus_size; i++) { - struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i]; + struct kvm_s2_mmu *mmu = kvm->arch.nested_mmus[i]; if (kvm_s2_mmu_valid(mmu)) kvm_stage2_unmap_range(mmu, 0, kvm_phys_size(mmu), may_block); } - kvm_invalidate_vncr_ipa(kvm, 0, BIT(kvm->arch.mmu.pgt->ia_bits)); + kvm_invalidate_vncr_ipa_all(kvm); } void kvm_nested_s2_flush(struct kvm *kvm) @@ -1304,7 +1318,7 @@ void kvm_nested_s2_flush(struct kvm *kvm) return; for (i = 0; i < kvm->arch.nested_mmus_size; i++) { - struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i]; + struct kvm_s2_mmu *mmu = kvm->arch.nested_mmus[i]; if (kvm_s2_mmu_valid(mmu)) kvm_stage2_flush_range(mmu, 0, kvm_phys_size(mmu)); @@ -1313,17 +1327,12 @@ void kvm_nested_s2_flush(struct kvm *kvm) void kvm_arch_flush_shadow_all(struct kvm *kvm) { - int i; - - for (i = 0; i < kvm->arch.nested_mmus_size; i++) { - struct kvm_s2_mmu *mmu = &kvm->arch.nested_mmus[i]; + for (int i = 0; i < kvm->arch.nested_mmus_size; i++) { + struct kvm_s2_mmu *mmu = kvm->arch.nested_mmus[i]; if (!WARN_ON(atomic_read(&mmu->refcnt))) kvm_free_stage2_pgd(mmu); } - kvfree(kvm->arch.nested_mmus); - kvm->arch.nested_mmus = NULL; - kvm->arch.nested_mmus_size = 0; kvm_uninit_stage2_mmu(kvm); } diff --git a/arch/arm64/kvm/sys_regs.c b/arch/arm64/kvm/sys_regs.c index 797e888bf939..e566bbce7d22 100644 --- a/arch/arm64/kvm/sys_regs.c +++ b/arch/arm64/kvm/sys_regs.c @@ -889,6 +889,7 @@ static u64 *demux_wb_reg(struct kvm_vcpu *vcpu, const struct sys_reg_desc *rd) struct kvm_guest_debug_arch *dbg = &vcpu->arch.vcpu_debug_state; switch (rd->Op2) { + case 0b001: case 0b100: return &dbg->dbg_bvr[rd->CRm]; case 0b101: diff --git a/arch/arm64/kvm/vgic/vgic-its.c b/arch/arm64/kvm/vgic/vgic-its.c index ed281fbf008b..7dc5ef5a6e33 100644 --- a/arch/arm64/kvm/vgic/vgic-its.c +++ b/arch/arm64/kvm/vgic/vgic-its.c @@ -1658,7 +1658,7 @@ static void vgic_mmio_write_its_baser(struct kvm *kvm, unsigned long val) { const struct vgic_its_abi *abi = vgic_its_get_abi(its); - u64 entry_size, table_type; + u64 old, entry_size, table_type; u64 reg, *regptr, clearbits = 0; /* When GITS_CTLR.Enable is 1, we ignore write accesses. */ @@ -1681,7 +1681,9 @@ static void vgic_mmio_write_its_baser(struct kvm *kvm, return; } - reg = update_64bit_reg(*regptr, addr & 7, len, val); + old = *regptr; + + reg = update_64bit_reg(old, addr & 7, len, val); reg &= ~GITS_BASER_RO_MASK; reg &= ~clearbits; @@ -1691,7 +1693,8 @@ static void vgic_mmio_write_its_baser(struct kvm *kvm, *regptr = reg; - if (!(reg & GITS_BASER_VALID)) { + /* The ITS driver rewrites an unchanged GITS_BASER on resume. */ + if (reg != old) { /* Take the its_lock to prevent a race with a save/restore */ mutex_lock(&its->its_lock); switch (table_type) { @@ -1702,6 +1705,8 @@ static void vgic_mmio_write_its_baser(struct kvm *kvm, vgic_its_free_collection_list(kvm, its); break; } + /* A concurrent injection may have cached a translation. */ + vgic_its_invalidate_cache(its); mutex_unlock(&its->its_lock); } } @@ -2019,18 +2024,22 @@ static int vgic_its_attr_regs_access(struct kvm_device *dev, return ret; } -static u32 compute_next_devid_offset(struct list_head *h, +static u32 compute_next_devid_offset(struct vgic_its *its, u64 baser, struct its_device *dev) { - struct its_device *next; - u32 next_offset; + struct its_device *next = dev; - if (list_is_last(&dev->dev_list, h)) - return 0; - next = list_next_entry(dev, dev_list); - next_offset = next->device_id - dev->device_id; + /* + * Point at the next device vgic_its_save_device_tables() saves. It + * sorts device_list first, so the subtraction cannot underflow. + */ + list_for_each_entry_continue(next, &its->device_list, dev_list) { + if (vgic_its_check_id(its, baser, next->device_id, NULL)) + return min_t(u32, next->device_id - dev->device_id, + VITS_DTE_MAX_DEVID_OFFSET); + } - return min_t(u32, next_offset, VITS_DTE_MAX_DEVID_OFFSET); + return 0; } static u32 compute_next_eventid_offset(struct list_head *h, struct its_ite *ite) @@ -2270,17 +2279,18 @@ static int vgic_its_restore_itt(struct vgic_its *its, struct its_device *dev) * vgic_its_save_dte - Save a device table entry at a given GPA * * @its: ITS handle + * @baser: GITS_BASER the caller is saving against * @dev: ITS device * @ptr: GPA */ -static int vgic_its_save_dte(struct vgic_its *its, struct its_device *dev, - gpa_t ptr) +static int vgic_its_save_dte(struct vgic_its *its, u64 baser, + struct its_device *dev, gpa_t ptr) { u64 val, itt_addr_field; u32 next_offset; itt_addr_field = dev->itt_addr >> 8; - next_offset = compute_next_devid_offset(&its->device_list, dev); + next_offset = compute_next_devid_offset(its, baser, dev); val = (1ULL << KVM_ITS_DTE_VALID_SHIFT | ((u64)next_offset << KVM_ITS_DTE_NEXT_SHIFT) | (itt_addr_field << KVM_ITS_DTE_ITTADDR_SHIFT) | @@ -2379,15 +2389,16 @@ static int vgic_its_save_device_tables(struct vgic_its *its) int ret; gpa_t eaddr; + /* Don't fail a save that userspace must be able to issue. */ if (!vgic_its_check_id(its, baser, dev->device_id, &eaddr)) - return -EINVAL; + continue; ret = vgic_its_save_itt(its, dev); if (ret) return ret; - ret = vgic_its_save_dte(its, dev, eaddr); + ret = vgic_its_save_dte(its, baser, dev, eaddr); if (ret) return ret; } diff --git a/arch/arm64/net/bpf_jit_comp.c b/arch/arm64/net/bpf_jit_comp.c index d4e62484baea..8c0c9b0fda16 100644 --- a/arch/arm64/net/bpf_jit_comp.c +++ b/arch/arm64/net/bpf_jit_comp.c @@ -600,6 +600,8 @@ static int build_prologue(struct jit_ctx *ctx, bool ebpf_from_cbpf) * 12 registers are on the stack */ emit(A64_SUB_I(1, A64_SP, A64_FP, 96), ctx); + /* The callback may use its own BPF stack, set up fp for it. */ + ctx->fp_used = true; } /* Stack must be multiples of 16B */ diff --git a/arch/mips/net/bpf_jit_comp32.c b/arch/mips/net/bpf_jit_comp32.c index 40a878b672f5..15a2a153dc87 100644 --- a/arch/mips/net/bpf_jit_comp32.c +++ b/arch/mips/net/bpf_jit_comp32.c @@ -1111,7 +1111,7 @@ static void emit_jmp_i64(struct jit_context *ctx, emit(ctx, xor, tmp, lo(dst), tmp); } if (imm < 0) { /* Compare sign extension */ - emit(ctx, addu, MIPS_R_T9, hi(dst), 1); + emit(ctx, addiu, MIPS_R_T9, hi(dst), 1); emit(ctx, or, tmp, tmp, MIPS_R_T9); } else { /* Compare zero extension */ emit(ctx, or, tmp, tmp, hi(dst)); diff --git a/arch/mips/net/bpf_jit_comp64.c b/arch/mips/net/bpf_jit_comp64.c index fa7e9aa37f49..6681ccac9dd9 100644 --- a/arch/mips/net/bpf_jit_comp64.c +++ b/arch/mips/net/bpf_jit_comp64.c @@ -305,8 +305,7 @@ static void emit_bswap_r64(struct jit_context *ctx, u8 dst, u32 width) case 16: emit_sext(ctx, dst, dst); emit_bswap_r(ctx, dst, width); - if (cpu_has_mips64r2 || cpu_has_mips64r6) - emit_zext(ctx, dst); + emit_zext(ctx, dst); break; } clobber_reg(ctx, dst); diff --git a/arch/parisc/include/asm/thread_info.h b/arch/parisc/include/asm/thread_info.h index b283738bb6da..249370706242 100644 --- a/arch/parisc/include/asm/thread_info.h +++ b/arch/parisc/include/asm/thread_info.h @@ -24,7 +24,7 @@ struct thread_info { /* thread information allocation */ -#ifdef CONFIG_IRQSTACKS +#if defined(CONFIG_IRQSTACKS) && !defined(CONFIG_64BIT) #define THREAD_SIZE_ORDER 2 /* PA-RISC requires at least 16k stack */ #else #define THREAD_SIZE_ORDER 3 /* PA-RISC requires at least 32k stack */ diff --git a/arch/parisc/kernel/setup.c b/arch/parisc/kernel/setup.c index d3e17a7a8901..4d3015a411d6 100644 --- a/arch/parisc/kernel/setup.c +++ b/arch/parisc/kernel/setup.c @@ -132,6 +132,18 @@ void __init setup_arch(char **cmdline_p) parisc_cache_init(); paging_init(); + /* + * Parse early parameters before mm_core_init_early() runs. + * Several early_param() handlers only record data that is consumed + * from there - for example hugepages=, hugepagesz=, + * default_hugepagesz=, hugetlb_cma= and hugetlb_free_vmemmap= - so + * the generic parse_early_param() call in start_kernel() is too late + * for them. jump_label_init() must come first, since early param + * handlers may enable or disable static keys. + */ + jump_label_init(); + parse_early_param(); + #ifdef CONFIG_PA11 dma_ops_init(); #endif diff --git a/arch/riscv/include/asm/csr.h b/arch/riscv/include/asm/csr.h index 31b8988f4488..72ff259154e9 100644 --- a/arch/riscv/include/asm/csr.h +++ b/arch/riscv/include/asm/csr.h @@ -183,12 +183,24 @@ #define HGATP_MODE_SHIFT HGATP32_MODE_SHIFT #endif -/* VSIP & HVIP relation */ +/* + * VSIP & HVIP relation + * + * The bit positions are same between VSIP and HVIP for interrupt + * numbers 13-63, where there's a shift for the SSI, STI and SEI. + */ #define VSIP_TO_HVIP_SHIFT (IRQ_VS_SOFT - IRQ_S_SOFT) -#define VSIP_VALID_MASK ((_AC(1, UL) << IRQ_S_SOFT) | \ +#define VSIP_BIAS_MASK ((_AC(1, UL) << IRQ_S_SOFT) | \ (_AC(1, UL) << IRQ_S_TIMER) | \ - (_AC(1, UL) << IRQ_S_EXT) | \ - (_AC(1, UL) << IRQ_PMU_OVF)) + (_AC(1, UL) << IRQ_S_EXT)) +#define VSIP_NO_BIAS_MASK (_AC(1, UL) << IRQ_PMU_OVF) +#define VSIP_VALID_MASK (VSIP_BIAS_MASK | VSIP_NO_BIAS_MASK) +#define vsip_to_hvip(_vsip) ((((_vsip) & VSIP_BIAS_MASK) << \ + VSIP_TO_HVIP_SHIFT) | \ + ((_vsip) & VSIP_NO_BIAS_MASK)) +#define hvip_to_vsip(_hvip) ((((_hvip) >> VSIP_TO_HVIP_SHIFT) & \ + VSIP_BIAS_MASK) | \ + ((_hvip) & VSIP_NO_BIAS_MASK)) /* AIA CSR bits */ #define TOPI_IID_SHIFT 16 diff --git a/arch/riscv/kvm/aia_imsic.c b/arch/riscv/kvm/aia_imsic.c index d38f5de0834c..74f520e572b1 100644 --- a/arch/riscv/kvm/aia_imsic.c +++ b/arch/riscv/kvm/aia_imsic.c @@ -969,9 +969,14 @@ int kvm_riscv_aia_imsic_rw_attr(struct kvm *kvm, unsigned long type, if (!vcpu) return -ENODEV; + if (mutex_lock_killable(&vcpu->mutex)) + return -EINTR; + imsic = vcpu->arch.aia_context.imsic_state; - if (!imsic) - return -ENODEV; + if (!imsic) { + rc = -ENODEV; + goto out_unlock; + } isel = KVM_DEV_RISCV_AIA_IMSIC_GET_ISEL(type); read_lock_irqsave(&imsic->vsfile_lock, flags); @@ -995,6 +1000,8 @@ int kvm_riscv_aia_imsic_rw_attr(struct kvm *kvm, unsigned long type, rc = imsic_vsfile_rw(vsfile_hgei, vsfile_cpu, imsic->nr_eix, isel, write, val); +out_unlock: + mutex_unlock(&vcpu->mutex); return rc; } diff --git a/arch/riscv/kvm/mmu.c b/arch/riscv/kvm/mmu.c index 8a0aa5e0e216..59e7297dea4a 100644 --- a/arch/riscv/kvm/mmu.c +++ b/arch/riscv/kvm/mmu.c @@ -540,6 +540,7 @@ int kvm_riscv_mmu_map(struct kvm_vcpu *vcpu, struct kvm_memory_slot *memslot, int ret; kvm_pfn_t hfn; bool is_hugetlb; + bool unused = false; bool writable; short vma_pageshift; gfn_t gfn = gpa >> PAGE_SHIFT; @@ -620,6 +621,8 @@ int kvm_riscv_mmu_map(struct kvm_vcpu *vcpu, struct kvm_memory_slot *memslot, vma_pageshift, current); return 0; } + if (is_sigpending_pfn(hfn)) + return -EINTR; if (is_error_noslot_pfn(hfn)) return -EFAULT; @@ -632,8 +635,10 @@ int kvm_riscv_mmu_map(struct kvm_vcpu *vcpu, struct kvm_memory_slot *memslot, write_lock(&kvm->mmu_lock); - if (mmu_invalidate_retry(kvm, mmu_seq)) + if (mmu_invalidate_retry(kvm, mmu_seq)) { + unused = true; goto out_unlock; + } /* * Check if we are backed by a THP and thus use block mapping if @@ -656,7 +661,8 @@ int kvm_riscv_mmu_map(struct kvm_vcpu *vcpu, struct kvm_memory_slot *memslot, kvm_err("Failed to map in G-stage\n"); out_unlock: - kvm_release_faultin_page(kvm, page, ret && ret != -EEXIST, writable); + kvm_release_faultin_page(kvm, page, + unused || (ret && ret != -EEXIST), writable); write_unlock(&kvm->mmu_lock); return ret; } diff --git a/arch/riscv/kvm/nacl.c b/arch/riscv/kvm/nacl.c index 6f9f8963e9dd..0a2a50c6ce03 100644 --- a/arch/riscv/kvm/nacl.c +++ b/arch/riscv/kvm/nacl.c @@ -42,12 +42,24 @@ void __kvm_riscv_nacl_hfence(void *shmem, } } - entp = shmem + SBI_NACL_SHMEM_HFENCE_ENTRY_CONFIG(i); - *entp = cpu_to_lelong(control); + /* + * Per SBI v3.0 section 15.1.2, the Page_Number and Page_Count + * words must be updated before the Config word with its Pending + * bit set. WRITE_ONCE() stops the compiler from reordering the + * stores and smp_wmb() makes the parameter words globally + * visible to the SBI implementation (or NACL hardware) before + * the Pending bit is set. + */ entp = shmem + SBI_NACL_SHMEM_HFENCE_ENTRY_PNUM(i); - *entp = cpu_to_lelong(page_num); + WRITE_ONCE(*entp, cpu_to_lelong(page_num)); entp = shmem + SBI_NACL_SHMEM_HFENCE_ENTRY_PCOUNT(i); - *entp = cpu_to_lelong(page_count); + WRITE_ONCE(*entp, cpu_to_lelong(page_count)); + + /* Ensure the parameter words are visible before the Pending bit */ + smp_wmb(); + + entp = shmem + SBI_NACL_SHMEM_HFENCE_ENTRY_CONFIG(i); + WRITE_ONCE(*entp, cpu_to_lelong(control)); } int kvm_riscv_nacl_enable(void) diff --git a/arch/riscv/kvm/vcpu.c b/arch/riscv/kvm/vcpu.c index 977e36ab83d3..468918309dff 100644 --- a/arch/riscv/kvm/vcpu.c +++ b/arch/riscv/kvm/vcpu.c @@ -475,8 +475,7 @@ bool kvm_riscv_vcpu_has_interrupts(struct kvm_vcpu *vcpu, u64 mask) bool ret; raw_spin_lock_irqsave(&vcpu->arch.irqs_pending_lock, flags); - ie = ((vcpu->arch.guest_csr.vsie & VSIP_VALID_MASK) - << VSIP_TO_HVIP_SHIFT) & (unsigned long)mask; + ie = vsip_to_hvip(vcpu->arch.guest_csr.vsie) & (unsigned long)mask; ie |= vcpu->arch.guest_csr.vsie & ~IRQ_LOCAL_MASK & (unsigned long)mask; ret = vcpu->arch.irqs_pending[0] & ie; diff --git a/arch/riscv/kvm/vcpu_exit.c b/arch/riscv/kvm/vcpu_exit.c index 6c8530b9f29e..8092200c0f47 100644 --- a/arch/riscv/kvm/vcpu_exit.c +++ b/arch/riscv/kvm/vcpu_exit.c @@ -267,7 +267,7 @@ int kvm_riscv_vcpu_exit(struct kvm_vcpu *vcpu, struct kvm_run *run, } /* Print details in-case of error */ - if (ret < 0) { + if (ret < 0 && ret != -EINTR) { kvm_err("VCPU exit error %d\n", ret); kvm_err("SEPC=0x%lx SSTATUS=0x%lx HSTATUS=0x%lx\n", vcpu->arch.guest_context.sepc, diff --git a/arch/riscv/kvm/vcpu_onereg.c b/arch/riscv/kvm/vcpu_onereg.c index 99b9107b1ac1..9fe829eed178 100644 --- a/arch/riscv/kvm/vcpu_onereg.c +++ b/arch/riscv/kvm/vcpu_onereg.c @@ -272,7 +272,7 @@ static int kvm_riscv_vcpu_general_get_csr(struct kvm_vcpu *vcpu, if (reg_num == KVM_REG_RISCV_CSR_REG(sip)) { kvm_riscv_vcpu_flush_interrupts(vcpu); - *out_val = (csr->hvip >> VSIP_TO_HVIP_SHIFT) & VSIP_VALID_MASK; + *out_val = hvip_to_vsip(csr->hvip); *out_val |= csr->hvip & ~IRQ_LOCAL_MASK; } else *out_val = ((unsigned long *)csr)[reg_num]; @@ -293,10 +293,8 @@ static int kvm_riscv_vcpu_general_set_csr(struct kvm_vcpu *vcpu, reg_num = array_index_nospec(reg_num, regs_max); - if (reg_num == KVM_REG_RISCV_CSR_REG(sip)) { - reg_val &= VSIP_VALID_MASK; - reg_val <<= VSIP_TO_HVIP_SHIFT; - } + if (reg_num == KVM_REG_RISCV_CSR_REG(sip)) + reg_val = vsip_to_hvip(reg_val); ((unsigned long *)csr)[reg_num] = reg_val; diff --git a/arch/riscv/kvm/vcpu_pmu.c b/arch/riscv/kvm/vcpu_pmu.c index a7f948410d53..73d76fd19896 100644 --- a/arch/riscv/kvm/vcpu_pmu.c +++ b/arch/riscv/kvm/vcpu_pmu.c @@ -270,12 +270,13 @@ static int pmu_ctr_read(struct kvm_vcpu *vcpu, unsigned long cidx, return -EINVAL; pmc->counter_val = kvpmu->fw_event[fevent_code].value; + *out_val = pmc->counter_val; } else if (pmc->perf_event) { - pmc->counter_val += perf_event_read_value(pmc->perf_event, &enabled, &running); + *out_val = pmc->counter_val + + perf_event_read_value(pmc->perf_event, &enabled, &running); } else { return -EINVAL; } - *out_val = pmc->counter_val; return 0; } @@ -645,7 +646,6 @@ int kvm_riscv_vcpu_pmu_ctr_stop(struct kvm_vcpu *vcpu, unsigned long ctr_base, { struct kvm_pmu *kvpmu = vcpu_to_pmu(vcpu); int i, pmc_index, sbiret = 0; - u64 enabled, running; struct kvm_pmc *pmc; int fevent_code; bool snap_flag_set = flags & SBI_PMU_STOP_FLAG_TAKE_SNAPSHOT; @@ -675,14 +675,19 @@ int kvm_riscv_vcpu_pmu_ctr_stop(struct kvm_vcpu *vcpu, unsigned long ctr_base, goto out; } - if (!kvpmu->fw_event[fevent_code].started) + if (!kvpmu->fw_event[fevent_code].started) { sbiret = SBI_ERR_ALREADY_STOPPED; - - kvpmu->fw_event[fevent_code].started = false; + } else { + kvpmu->fw_event[fevent_code].started = false; + pmc->counter_val = kvpmu->fw_event[fevent_code].value; + } } else if (pmc->perf_event) { if (pmc->started) { - /* Stop counting the counter */ - perf_event_disable(pmc->perf_event); + /* + * Stop the counter and fold the live count into counter_val. + * Reset the event value to avoid redundant accumulation. + */ + pmc->counter_val += perf_event_pause(pmc->perf_event, true); pmc->started = false; } else { sbiret = SBI_ERR_ALREADY_STOPPED; @@ -696,11 +701,6 @@ int kvm_riscv_vcpu_pmu_ctr_stop(struct kvm_vcpu *vcpu, unsigned long ctr_base, } if (snap_flag_set && !sbiret) { - if (pmc->cinfo.type == SBI_PMU_CTR_TYPE_FW) - pmc->counter_val = kvpmu->fw_event[fevent_code].value; - else if (pmc->perf_event) - pmc->counter_val += perf_event_read_value(pmc->perf_event, - &enabled, &running); /* * The counter and overflow indices in the snapshot region are w.r.to * cbase. Modify the set bit in the counter mask instead of the pmc_index @@ -727,9 +727,10 @@ int kvm_riscv_vcpu_pmu_ctr_stop(struct kvm_vcpu *vcpu, unsigned long ctr_base, } } - if (shmem_needs_update) - kvm_vcpu_write_guest(vcpu, kvpmu->snapshot_addr, kvpmu->sdata, - sizeof(struct riscv_pmu_snapshot_data)); + if (shmem_needs_update && + kvm_vcpu_write_guest(vcpu, kvpmu->snapshot_addr, kvpmu->sdata, + sizeof(struct riscv_pmu_snapshot_data))) + sbiret = SBI_ERR_FAILURE; out: retdata->err_val = sbiret; diff --git a/arch/riscv/kvm/vcpu_sbi_hsm.c b/arch/riscv/kvm/vcpu_sbi_hsm.c index f26207f84bab..06a15629c26b 100644 --- a/arch/riscv/kvm/vcpu_sbi_hsm.c +++ b/arch/riscv/kvm/vcpu_sbi_hsm.c @@ -95,9 +95,9 @@ static int kvm_sbi_ext_hsm_handler(struct kvm_vcpu *vcpu, struct kvm_run *run, ret = kvm_sbi_hsm_vcpu_get_status(vcpu); if (ret >= 0) { retdata->out_val = ret; - retdata->err_val = 0; + ret = 0; } - return 0; + break; case SBI_EXT_HSM_HART_SUSPEND: switch (lower_32_bits(cp->a0)) { case SBI_HSM_SUSPEND_RET_DEFAULT: diff --git a/arch/riscv/kvm/vcpu_timer.c b/arch/riscv/kvm/vcpu_timer.c index ae53133c7ab0..a2cd277a4059 100644 --- a/arch/riscv/kvm/vcpu_timer.c +++ b/arch/riscv/kvm/vcpu_timer.c @@ -61,10 +61,13 @@ static enum hrtimer_restart kvm_riscv_vcpu_hrtimer_expired(struct hrtimer *h) static int kvm_riscv_vcpu_timer_cancel(struct kvm_vcpu_timer *t) { - if (!t->init_done || !t->next_set) + if (!t->init_done) return -EINVAL; hrtimer_cancel(&t->hrt); + + if (!t->next_set) + return -EINVAL; t->next_set = false; return 0; diff --git a/arch/s390/crypto/hmac_s390.c b/arch/s390/crypto/hmac_s390.c index f8cd09f341d4..445fa7bbd958 100644 --- a/arch/s390/crypto/hmac_s390.c +++ b/arch/s390/crypto/hmac_s390.c @@ -150,7 +150,10 @@ static int hash_data(const u8 *in, unsigned int inlen, #undef PARAM_INIT - cpacf_klmd(func, ¶m, in, inlen); + if (final) + cpacf_klmd(func, ¶m, in, inlen); + else + cpacf_kimd(func, ¶m, in, inlen); memcpy(digest, ¶m, digestsize); diff --git a/arch/s390/include/asm/debug.h b/arch/s390/include/asm/debug.h index 39d484c59774..ad438d6352c8 100644 --- a/arch/s390/include/asm/debug.h +++ b/arch/s390/include/asm/debug.h @@ -460,7 +460,11 @@ static int VNAME(var, active_entries)[EARLY_AREAS] __initdata #define __REGISTER_STATIC_DEBUG_INFO(var, name, pages, areas, view) \ static int __init VNAME(var, reg)(void) \ { \ - debug_register_static(&var, (pages), (areas)); \ + int rc; \ + \ + rc = debug_register_static(&var, (pages), (areas)); \ + if (rc) \ + return rc; \ debug_register_view(&var, (view)); \ return 0; \ } \ @@ -493,7 +497,7 @@ static debug_info_t __refdata var = \ static debug_info_t __used __section(".s390dbf_info") *VNAME(var, info) = &var; \ __REGISTER_STATIC_DEBUG_INFO(var, name, pages, nr_areas, view) -void debug_register_static(debug_info_t *id, int pages_per_area, int nr_areas); +int debug_register_static(debug_info_t *id, int pages_per_area, int nr_areas); #endif /* MODULE */ diff --git a/arch/s390/kernel/debug.c b/arch/s390/kernel/debug.c index e06abf1dbc21..71c966f4c5d0 100644 --- a/arch/s390/kernel/debug.c +++ b/arch/s390/kernel/debug.c @@ -432,7 +432,8 @@ static debug_info_t *debug_info_copy(debug_info_t *in, int mode) debug_info_free(rc); } while (1); - if (mode == NO_AREAS) + /* debug_register_static() failure leaves areas NULL, bounds intact */ + if (mode == NO_AREAS || !in->areas) goto out; for (i = 0; i < in->nr_areas; i++) { @@ -948,8 +949,12 @@ EXPORT_SYMBOL(debug_register); * * Note: This function is called automatically via an initcall generated by * DEFINE_STATIC_DEBUG_INFO. + * + * Return: + * - 0 on success + * - negative error code on failure */ -void debug_register_static(debug_info_t *id, int pages_per_area, int nr_areas) +int debug_register_static(debug_info_t *id, int pages_per_area, int nr_areas) { unsigned long flags; debug_info_t *copy; @@ -957,7 +962,7 @@ void debug_register_static(debug_info_t *id, int pages_per_area, int nr_areas) if (!initialized) { pr_err("Tried to register debug feature %s too early\n", id->name); - return; + return -EINVAL; } debug_get_param(id->name, &id->level, &pages_per_area); @@ -973,7 +978,7 @@ void debug_register_static(debug_info_t *id, int pages_per_area, int nr_areas) id->active_entries = NULL; raw_spin_unlock_irqrestore(&id->lock, flags); - return; + return -ENOMEM; } /* Replace static trace area with dynamic copy. */ @@ -991,6 +996,8 @@ void debug_register_static(debug_info_t *id, int pages_per_area, int nr_areas) mutex_lock(&debug_mutex); _debug_register(id); mutex_unlock(&debug_mutex); + + return 0; } /* Remove debugfs entries. */ diff --git a/arch/s390/kernel/uv.c b/arch/s390/kernel/uv.c index a284f98d9716..e70acad09cd5 100644 --- a/arch/s390/kernel/uv.c +++ b/arch/s390/kernel/uv.c @@ -760,7 +760,7 @@ static int find_secret_in_page(const u8 secret_id[UV_SECRET_ID_LEN], { u16 i; - for (i = 0; i < list->total_num_secrets; i++) { + for (i = 0; i < list->num_secr_stored; i++) { if (memcmp(secret_id, list->secrets[i].id, UV_SECRET_ID_LEN) == 0) { *secret = list->secrets[i].hdr; return 0; @@ -781,11 +781,14 @@ int uv_find_secret(const u8 secret_id[UV_SECRET_ID_LEN], struct uv_secret_list *list, struct uv_secret_list_item_hdr *secret) { - u16 start_idx = 0; + u16 start_idx; u16 list_rc; int ret; + list->next_secret_idx = 0; + do { + start_idx = list->next_secret_idx; uv_list_secrets(list, start_idx, &list_rc, NULL); if (list_rc != UVC_RC_EXECUTED && list_rc != UVC_RC_MORE_DATA) { if (list_rc == UVC_RC_INV_CMD) @@ -796,7 +799,6 @@ int uv_find_secret(const u8 secret_id[UV_SECRET_ID_LEN], ret = find_secret_in_page(secret_id, list, secret); if (ret == 0) return ret; - start_idx = list->next_secret_idx; } while (list_rc == UVC_RC_MORE_DATA && start_idx < list->next_secret_idx); return -ENOENT; diff --git a/arch/s390/kvm/dat.c b/arch/s390/kvm/dat.c index 3f2d6e8902d7..a21cb3975e99 100644 --- a/arch/s390/kvm/dat.c +++ b/arch/s390/kvm/dat.c @@ -620,17 +620,20 @@ int dat_get_storage_key(union asce asce, gfn_t gfn, union skey *skey) union pte *ptep; int rc; +again: skey->skey = 0; rc = dat_entry_walk(NULL, gfn, asce, DAT_WALK_ANY, TABLE_TYPE_PAGE_TABLE, &crstep, &ptep); if (rc) return rc; if (!ptep) { - union crste crste; + union crste crste = READ_ONCE(*crstep); - crste = READ_ONCE(*crstep); - if (!crste.h.fc || !crste.s.fc1.pr) + if (!crste_leaf(crste) && !crste.h.i) + goto again; + if (!crste.s.fc1.pr) return 0; + skey->skey = page_get_storage_key(large_crste_to_phys(crste, gfn)); return 0; } @@ -661,13 +664,20 @@ int dat_set_storage_key(struct kvm_s390_mmu_cache *mc, union asce asce, gfn_t gf union pte *ptep; int rc; +again: rc = dat_entry_walk(mc, gfn, asce, DAT_WALK_LEAF_ALLOC, TABLE_TYPE_PAGE_TABLE, &crstep, &ptep); if (rc) return rc; if (!ptep) { - page_set_storage_key(large_crste_to_phys(*crstep, gfn), skey.skey, !nq); + union crste crste = READ_ONCE(*crstep); + + /* A large page has been split concurrently, try again */ + if (!crste_leaf(crste)) + goto again; + + page_set_storage_key(large_crste_to_phys(crste, gfn), skey.skey, !nq); return 0; } @@ -717,14 +727,24 @@ int dat_cond_set_storage_key(struct kvm_s390_mmu_cache *mmc, union asce asce, gf union pte *ptep; int rc; +again: rc = dat_entry_walk(mmc, gfn, asce, DAT_WALK_LEAF_ALLOC, TABLE_TYPE_PAGE_TABLE, &crstep, &ptep); if (rc) return rc; - if (!ptep) - return page_cond_set_storage_key(large_crste_to_phys(*crstep, gfn), skey, oldkey, + if (!ptep) { + union crste crste = READ_ONCE(*crstep); + + /* A large page has been split concurrently, try again */ + if (!crste_leaf(crste)) + goto again; + if (!oldkey) + oldkey = &prev; + + return page_cond_set_storage_key(large_crste_to_phys(crste, gfn), skey, oldkey, nq, mr, mc); + } old = pgste_get_lock(ptep); pgste = old; @@ -763,7 +783,7 @@ int dat_reset_reference_bit(union asce asce, gfn_t gfn, union skey *skey) int rc; skey->skey = 0; - +again: rc = dat_entry_walk(NULL, gfn, asce, DAT_WALK_ANY, TABLE_TYPE_PAGE_TABLE, &crstep, &ptep); if (rc) return rc; @@ -771,9 +791,12 @@ int dat_reset_reference_bit(union asce asce, gfn_t gfn, union skey *skey) if (!ptep) { union crste crste = READ_ONCE(*crstep); - if (!crste.h.fc || !crste.s.fc1.pr) + /* A large page has been split concurrently, try again */ + if (!crste_leaf(crste) && !crste.h.i) + goto again; + if (!crste.s.fc1.pr) return 0; - skey->skey = page_reset_referenced(large_crste_to_phys(*crstep, gfn)) << 1; + skey->skey = page_reset_referenced(large_crste_to_phys(crste, gfn)) << 1; return 0; } old = pgste_get_lock(ptep); diff --git a/arch/s390/kvm/gaccess.c b/arch/s390/kvm/gaccess.c index 36102b2727fb..7c1f614ec314 100644 --- a/arch/s390/kvm/gaccess.c +++ b/arch/s390/kvm/gaccess.c @@ -1593,12 +1593,25 @@ static inline int ___gaccess_shadow_fault(struct kvm_vcpu *vcpu, struct gmap *sg parent = READ_ONCE(sg->parent); if (!parent) return -EAGAIN; +retry: scoped_guard(spinlock, &parent->children_lock) { if (READ_ONCE(sg->parent) != parent) return -EAGAIN; sg->invalidated = false; rc = _gaccess_do_shadow(vcpu->arch.mc, sg, saddr, walk); } + if (rc == -ENOENT) { + struct kvm_memory_slot *slot; + struct guest_fault *entries; + + entries = get_entries(walk); + slot = kvm_vcpu_gfn_to_memslot(vcpu, entries[LEVEL_MEM].gfn); + if (!slot) + return PGM_ADDRESSING; + rc = gmap_link(vcpu->arch.mc, parent, entries + LEVEL_MEM, slot); + if (!rc) + goto retry; + } if (!rc) kvm_s390_release_faultin_array(vcpu->kvm, walk->raw_entries, false); return rc; diff --git a/arch/s390/kvm/gmap.c b/arch/s390/kvm/gmap.c index 8abb4f55b306..f11d4ecaef7b 100644 --- a/arch/s390/kvm/gmap.c +++ b/arch/s390/kvm/gmap.c @@ -991,11 +991,13 @@ static long _destroy_pages_pte(union pte *ptep, gfn_t gfn, gfn_t next, struct da static long _destroy_pages_crste(union crste *crstep, gfn_t gfn, gfn_t next, struct dat_walk *walk) { phys_addr_t origin, cur, end; + union crste crste; - if (!crstep->h.fc || !crstep->s.fc1.pr) + crste = READ_ONCE(*crstep); + if (!crste.h.fc || !crste.s.fc1.pr) return 0; - origin = crste_origin_large(*crstep); + origin = crste_origin_large(crste); cur = ((max(gfn, walk->start) - gfn) << PAGE_SHIFT) + origin; end = ((min(next, walk->end) - gfn) << PAGE_SHIFT) + origin; for ( ; cur < end; cur += PAGE_SIZE) diff --git a/arch/s390/kvm/interrupt.c b/arch/s390/kvm/interrupt.c index d33dd58f4a4f..80f4802934df 100644 --- a/arch/s390/kvm/interrupt.c +++ b/arch/s390/kvm/interrupt.c @@ -1550,23 +1550,21 @@ static int __inject_set_prefix(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq) } #define KVM_S390_STOP_SUPP_FLAGS (KVM_S390_STOP_FLAG_STORE_STATUS) -static int __inject_sigp_stop(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq) +static int __inject_sigp_stop(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq, bool *storestatus) { struct kvm_s390_local_interrupt *li = &vcpu->arch.local_int; struct kvm_s390_stop_info *stop = &li->irq.stop; - int rc = 0; vcpu->stat.inject_stop_signal++; trace_kvm_s390_inject_vcpu(vcpu->vcpu_id, KVM_S390_SIGP_STOP, 0, 0); if (irq->u.stop.flags & ~KVM_S390_STOP_SUPP_FLAGS) return -EINVAL; - if (is_vcpu_stopped(vcpu)) { - if (irq->u.stop.flags & KVM_S390_STOP_FLAG_STORE_STATUS) - rc = kvm_s390_store_status_unloaded(vcpu, - KVM_S390_STORE_STATUS_NOADDR); - return rc; + if (!(irq->u.stop.flags & KVM_S390_STOP_FLAG_STORE_STATUS)) + return 0; + *storestatus = true; + return -EWOULDBLOCK; } if (test_and_set_bit(IRQ_PEND_SIGP_STOP, &li->pending_irqs)) @@ -2102,7 +2100,7 @@ void kvm_s390_clear_stop_irq(struct kvm_vcpu *vcpu) spin_unlock(&li->lock); } -static int do_inject_vcpu(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq) +static int do_inject_vcpu(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq, bool *storestatus) { int rc; @@ -2114,7 +2112,7 @@ static int do_inject_vcpu(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq) rc = __inject_set_prefix(vcpu, irq); break; case KVM_S390_SIGP_STOP: - rc = __inject_sigp_stop(vcpu, irq); + rc = __inject_sigp_stop(vcpu, irq, storestatus); break; case KVM_S390_RESTART: rc = __inject_sigp_restart(vcpu); @@ -2150,11 +2148,16 @@ static int do_inject_vcpu(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq) int kvm_s390_inject_vcpu(struct kvm_vcpu *vcpu, struct kvm_s390_irq *irq) { struct kvm_s390_local_interrupt *li = &vcpu->arch.local_int; + bool storestatus = false; int rc; spin_lock(&li->lock); - rc = do_inject_vcpu(vcpu, irq); + rc = do_inject_vcpu(vcpu, irq, &storestatus); spin_unlock(&li->lock); + + if (rc == -EWOULDBLOCK && storestatus) + rc = kvm_s390_store_status_unloaded(vcpu, KVM_S390_STORE_STATUS_NOADDR); + if (!rc) kvm_s390_vcpu_wakeup(vcpu); return rc; @@ -2976,61 +2979,58 @@ static int adapter_indicators_set(struct kvm *kvm, struct s390_io_adapter *adapter, struct kvm_s390_adapter_int *adapter_int) { - unsigned long bit; - int summary_set, idx; struct s390_map_info *ind_info, *summary_info; - void *map; struct page *ind_page, *summary_page; - unsigned long flags; + unsigned long bit; + int summary_set; + void *map; ind_page = NULL; - spin_lock_irqsave(&adapter->maps_lock, flags); - ind_info = get_map_info(adapter, adapter_int->ind_addr); + scoped_guard(spinlock_irqsave, &adapter->maps_lock) { + ind_info = get_map_info(adapter, adapter_int->ind_addr); + if (ind_info) { + map = page_address(ind_info->page); + bit = get_ind_bit(ind_info->addr, adapter_int->ind_offset, adapter->swap); + set_bit(bit, map); + } + } if (!ind_info) { - spin_unlock_irqrestore(&adapter->maps_lock, flags); ind_page = pin_map_page(kvm, adapter_int->ind_addr, 0); if (!ind_page) return -1; - idx = srcu_read_lock(&kvm->srcu); map = page_address(ind_page); bit = get_ind_bit(adapter_int->ind_addr, adapter_int->ind_offset, adapter->swap); set_bit(bit, map); - mark_page_dirty(kvm, adapter_int->ind_gaddr >> PAGE_SHIFT); set_page_dirty_lock(ind_page); - srcu_read_unlock(&kvm->srcu, idx); unpin_user_page(ind_page); - } else { - map = page_address(ind_info->page); - bit = get_ind_bit(ind_info->addr, adapter_int->ind_offset, adapter->swap); - set_bit(bit, map); - spin_unlock_irqrestore(&adapter->maps_lock, flags); } + scoped_guard(srcu, &kvm->srcu) + mark_page_dirty(kvm, gpa_to_gfn(adapter_int->ind_gaddr)); - spin_lock_irqsave(&adapter->maps_lock, flags); - summary_info = get_map_info(adapter, adapter_int->summary_addr); + scoped_guard(spinlock_irqsave, &adapter->maps_lock) { + summary_info = get_map_info(adapter, adapter_int->summary_addr); + if (summary_info) { + map = page_address(summary_info->page); + bit = get_ind_bit(summary_info->addr, adapter_int->summary_offset, + adapter->swap); + summary_set = test_and_set_bit(bit, map); + } + } if (!summary_info) { - spin_unlock_irqrestore(&adapter->maps_lock, flags); summary_page = pin_map_page(kvm, adapter_int->summary_addr, 0); if (WARN_ON_ONCE(!summary_page)) return -1; - idx = srcu_read_lock(&kvm->srcu); map = page_address(summary_page); bit = get_ind_bit(adapter_int->summary_addr, adapter_int->summary_offset, adapter->swap); summary_set = test_and_set_bit(bit, map); - mark_page_dirty(kvm, adapter_int->summary_gaddr >> PAGE_SHIFT); set_page_dirty_lock(summary_page); - srcu_read_unlock(&kvm->srcu, idx); unpin_user_page(summary_page); - } else { - map = page_address(summary_info->page); - bit = get_ind_bit(summary_info->addr, adapter_int->summary_offset, - adapter->swap); - summary_set = test_and_set_bit(bit, map); - spin_unlock_irqrestore(&adapter->maps_lock, flags); } + scoped_guard(srcu, &kvm->srcu) + mark_page_dirty(kvm, gpa_to_gfn(adapter_int->summary_gaddr)); return summary_set ? 0 : 1; } @@ -3040,26 +3040,29 @@ static int adapter_indicators_set_fast(struct kvm *kvm, struct kvm_s390_adapter_int *adapter_int, int setbit) { + struct s390_map_info *ind_info, *summary_info; unsigned long bit; int summary_set; - struct s390_map_info *ind_info, *summary_info; void *map; - spin_lock(&adapter->maps_lock); + guard(srcu)(&kvm->srcu); + guard(spinlock)(&adapter->maps_lock); + ind_info = get_map_info(adapter, adapter_int->ind_addr); - if (!ind_info) { - spin_unlock(&adapter->maps_lock); + if (!ind_info) return -EWOULDBLOCK; - } + map = page_address(ind_info->page); bit = get_ind_bit(ind_info->addr, adapter_int->ind_offset, adapter->swap); - if (setbit) + if (setbit) { set_bit(bit, map); + mark_page_dirty(kvm, gpa_to_gfn(adapter_int->ind_gaddr)); + } + summary_info = get_map_info(adapter, adapter_int->summary_addr); - if (!summary_info) { - spin_unlock(&adapter->maps_lock); + if (!summary_info) return -EWOULDBLOCK; - } + map = page_address(summary_info->page); bit = get_ind_bit(summary_info->addr, adapter_int->summary_offset, adapter->swap); @@ -3069,7 +3072,8 @@ static int adapter_indicators_set_fast(struct kvm *kvm, summary_set = test_and_set_bit(bit, map); else summary_set = test_and_clear_bit(bit, map); - spin_unlock(&adapter->maps_lock); + mark_page_dirty(kvm, gpa_to_gfn(adapter_int->summary_gaddr)); + return summary_set ? 0 : 1; } @@ -3189,7 +3193,8 @@ int kvm_set_msi(struct kvm_kernel_irq_routing_entry *e, struct kvm *kvm, int kvm_s390_set_irq_state(struct kvm_vcpu *vcpu, void __user *irqstate, int len) { struct kvm_s390_local_interrupt *li = &vcpu->arch.local_int; - struct kvm_s390_irq *buf; + struct kvm_s390_irq *buf __free(kvfree) = NULL; + bool tmp, storestatus = false; int r = 0; int n; @@ -3197,31 +3202,33 @@ int kvm_s390_set_irq_state(struct kvm_vcpu *vcpu, void __user *irqstate, int len if (!buf) return -ENOMEM; - if (copy_from_user((void *) buf, irqstate, len)) { - r = -EFAULT; - goto out_free; - } + if (copy_from_user((void *)buf, irqstate, len)) + return -EFAULT; - /* - * Don't allow setting the interrupt state - * when there are already interrupts pending - */ - spin_lock(&li->lock); - if (li->pending_irqs) { - r = -EBUSY; - goto out_unlock; - } + scoped_guard(spinlock, &li->lock) { + /* + * Don't allow setting the interrupt state + * when there are already interrupts pending + */ + if (li->pending_irqs) + return -EBUSY; - for (n = 0; n < len / sizeof(*buf); n++) { - r = do_inject_vcpu(vcpu, &buf[n]); - if (r) - break; + for (n = 0; n < len / sizeof(*buf); n++) { + tmp = false; + r = do_inject_vcpu(vcpu, &buf[n], &tmp); + if (r == -EWOULDBLOCK && tmp) { + storestatus = true; + r = 0; + } + if (r) + break; + } + } + if (storestatus) { + scoped_guard(srcu, &vcpu->kvm->srcu) + n = kvm_s390_store_status_unloaded(vcpu, KVM_S390_STORE_STATUS_NOADDR); + return r ? r : n; } - -out_unlock: - spin_unlock(&li->lock); -out_free: - vfree(buf); return r; } diff --git a/arch/s390/pci/pci_event.c b/arch/s390/pci/pci_event.c index 839bd91c056e..c1d6ec9977a5 100644 --- a/arch/s390/pci/pci_event.c +++ b/arch/s390/pci/pci_event.c @@ -190,6 +190,7 @@ static pci_ers_result_t zpci_event_attempt_error_recovery(struct pci_dev *pdev) device_lock(&pdev->dev); if (pdev->error_state == pci_channel_io_perm_failure) { ers_res = PCI_ERS_RESULT_DISCONNECT; + status_str = "skipped (permanent failure)"; goto out_unlock; } pdev->error_state = pci_channel_io_frozen; @@ -256,8 +257,8 @@ static pci_ers_result_t zpci_event_attempt_error_recovery(struct pci_dev *pdev) driver->err_handler->resume(pdev); pci_uevent_ers(pdev, PCI_ERS_RESULT_RECOVERED); out_unlock: + zpci_report_status(zdev, pdev, "recovery", status_str); device_unlock(&pdev->dev); - zpci_report_status(zdev, "recovery", status_str); return ers_res; } @@ -320,8 +321,10 @@ static void __zpci_event_error(struct zpci_ccdf_err *ccdf) pr_err("%s: Event 0x%x reports an error for PCI function 0x%x\n", pdev ? pci_name(pdev) : "n/a", ccdf->pec, ccdf->fid); - if (!pdev) + if (!pdev) { + zpci_report_status(zdev, NULL, "error event", "no pdev bound"); goto no_pdev; + } switch (ccdf->pec) { case 0x002a: /* Error event concerns FMB */ diff --git a/arch/s390/pci/pci_report.c b/arch/s390/pci/pci_report.c index 7030f7052926..72ecabf7003c 100644 --- a/arch/s390/pci/pci_report.c +++ b/arch/s390/pci/pci_report.c @@ -87,9 +87,23 @@ static struct debug_view debug_log_view = { NULL }; +static ssize_t zpci_report_pdev(struct pci_dev *pdev, char *buf, size_t size) +{ + struct pci_driver *driver; + const char *start = buf; + char *end = buf + size; + + device_lock_assert(&pdev->dev); + buf += scnprintf(buf, end - buf, "state: %s\n", zpci_state_str(pdev->error_state)); + driver = to_pci_driver(pdev->dev.driver); + buf += scnprintf(buf, end - buf, "driver: %s\n", (driver) ? driver->name : "n/a"); + return buf - start; +} + /** * zpci_report_status - Report the status of operations on a PCI device - * @zdev: The PCI device for which to report status + * @zdev: The zPCI device for which to report status + * @pdev: The PCI device associated with the zdev if any, NULL otherwise * @operation: A string representing the operation reported * @status: A string representing the status of the operation * @@ -103,15 +117,14 @@ static struct debug_view debug_log_view = { * * Return: 0 on success an error code < 0 otherwise. */ -int zpci_report_status(struct zpci_dev *zdev, const char *operation, const char *status) +int zpci_report_status(struct zpci_dev *zdev, struct pci_dev *pdev, + const char *operation, const char *status) { struct zpci_report_error *report; - struct pci_driver *driver = NULL; - struct pci_dev *pdev = NULL; char *buf, *end; int ret; - if (!zdev || !zdev->zbus) + if (!zdev) return -ENODEV; /* Protected virtualization hosts get nothing from us */ @@ -121,18 +134,13 @@ int zpci_report_status(struct zpci_dev *zdev, const char *operation, const char report = (void *)get_zeroed_page(GFP_KERNEL); if (!report) return -ENOMEM; - if (zdev->zbus->bus) - pdev = pci_get_slot(zdev->zbus->bus, zdev->devfn); - if (pdev) - driver = to_pci_driver(pdev->dev.driver); buf = report->data.log_data; end = report->data.log_data + ZPCI_REPORT_DATA_SIZE; buf += scnprintf(buf, end - buf, "report: %s\n", operation); buf += scnprintf(buf, end - buf, "status: %s\n", status); - buf += scnprintf(buf, end - buf, "state: %s\n", - (pdev) ? zpci_state_str(pdev->error_state) : "n/a"); - buf += scnprintf(buf, end - buf, "driver: %s\n", (driver) ? driver->name : "n/a"); + if (pdev) + buf += zpci_report_pdev(pdev, buf, end - buf); ret = debug_dump(pci_debug_msg_id, &debug_log_view, buf, end - buf, true); if (ret < 0) pr_err("Reading PCI debug messages failed with code %d\n", ret); diff --git a/arch/s390/pci/pci_report.h b/arch/s390/pci/pci_report.h index e08003d51a97..dd7b0b05001c 100644 --- a/arch/s390/pci/pci_report.h +++ b/arch/s390/pci/pci_report.h @@ -8,9 +8,11 @@ */ #ifndef __S390_PCI_REPORT_H #define __S390_PCI_REPORT_H +#include struct zpci_dev; -int zpci_report_status(struct zpci_dev *zdev, const char *operation, const char *status); +int zpci_report_status(struct zpci_dev *zdev, struct pci_dev *pdev, + const char *operation, const char *status); #endif /* __S390_PCI_REPORT_H */ diff --git a/arch/x86/coco/sev/svsm.c b/arch/x86/coco/sev/svsm.c index 916d62cd17dc..2d493d55ff2e 100644 --- a/arch/x86/coco/sev/svsm.c +++ b/arch/x86/coco/sev/svsm.c @@ -74,6 +74,14 @@ int svsm_perform_call_protocol(struct svsm_call *call) flags = native_local_irq_save(); + /* + * 'caa' is a per-CPU variable. To avoid using a stale or incorrect + * 'caa' if the task is preempted or migrated to another CPU after it + * is fetched, always fetch 'caa' and then issue the SVSM call with + * interrupts disabled. This ensures the correct 'caa' is used. + */ + call->caa = svsm_get_caa(); + ghcb = __sev_get_ghcb(&state); do { @@ -321,7 +329,6 @@ int snp_svsm_vtpm_send_command(u8 *buffer) { struct svsm_call call = {}; - call.caa = svsm_get_caa(); call.rax = SVSM_VTPM_CALL(SVSM_VTPM_CMD); call.rcx = __pa(buffer); @@ -345,7 +352,6 @@ bool snp_svsm_vtpm_probe(void) if (!snp_vmpl) return false; - call.caa = svsm_get_caa(); call.rax = SVSM_VTPM_CALL(SVSM_VTPM_QUERY); if (svsm_perform_call_protocol(&call)) diff --git a/arch/x86/events/amd/brs.c b/arch/x86/events/amd/brs.c index dc564688f3d7..54b13faba116 100644 --- a/arch/x86/events/amd/brs.c +++ b/arch/x86/events/amd/brs.c @@ -343,11 +343,10 @@ void amd_brs_drain(void) if (!amd_brs_match_plm(event, from, to)) continue; - perf_clear_branch_entry_bitfields(br+nr); - - br[nr].from = from; - br[nr].to = to; - + br[nr] = (struct perf_branch_entry){ + .from = from, + .to = to, + }; nr++; } empty: diff --git a/arch/x86/events/amd/lbr.c b/arch/x86/events/amd/lbr.c index 9d9c961989d5..a55646fcb846 100644 --- a/arch/x86/events/amd/lbr.c +++ b/arch/x86/events/amd/lbr.c @@ -184,13 +184,6 @@ void amd_pmu_lbr_read(void) entry.to.split.reserved) continue; - perf_clear_branch_entry_bitfields(br + out); - - br[out].from = sign_ext_branch_ip(entry.from.split.ip); - br[out].to = sign_ext_branch_ip(entry.to.split.ip); - br[out].mispred = entry.from.split.mispredict; - br[out].predicted = !br[out].mispred; - /* * Set branch speculation information using the status of * the valid and spec bits. @@ -208,7 +201,14 @@ void amd_pmu_lbr_read(void) * speculative and took the correct path */ idx = (entry.to.split.valid << 1) | entry.to.split.spec; - br[out].spec = lbr_spec_map[idx]; + + br[out] = (struct perf_branch_entry){ + .from = sign_ext_branch_ip(entry.from.split.ip), + .to = sign_ext_branch_ip(entry.to.split.ip), + .mispred = entry.from.split.mispredict, + .predicted = !entry.from.split.mispredict, + .spec = lbr_spec_map[idx], + }; out++; } diff --git a/arch/x86/events/intel/core.c b/arch/x86/events/intel/core.c index 24c1f4059544..6437404c5c2c 100644 --- a/arch/x86/events/intel/core.c +++ b/arch/x86/events/intel/core.c @@ -506,6 +506,8 @@ static struct event_constraint intel_pnc_event_constraints[] = { INTEL_EVENT_CONSTRAINT(0xce, 0x1), INTEL_UEVENT_CONSTRAINT(0x01b1, 0x8), + INTEL_UEVENT_CONSTRAINT(0x01b2, 0xf), + INTEL_UEVENT_CONSTRAINT(0x02b2, 0xf), INTEL_UEVENT_CONSTRAINT(0x0847, 0xf), INTEL_UEVENT_CONSTRAINT(0x0446, 0xf), INTEL_UEVENT_CONSTRAINT(0x0846, 0xf), @@ -5294,12 +5296,15 @@ static struct perf_guest_switch_msr *intel_guest_get_msrs(int *nr, void *data) struct kvm_pmu *kvm_pmu = (struct kvm_pmu *)data; u64 intel_ctrl = hybrid(cpuc->pmu, intel_ctrl); u64 pebs_mask = cpuc->pebs_enabled & x86_pmu.pebs_capable; - int global_ctrl, pebs_enable; + u64 guest_pebs_mask; + int global_ctrl; /* * In addition to obeying exclude_guest/exclude_host, remove bits being * used for PEBS when running a guest, because PEBS writes to virtual - * addresses (not physical addresses). + * addresses (not physical addresses). If the guest wants to utilize + * PEBS, and PEBS can be safely enabled in the guest, bits for the guest's + * PEBS-enabled counters will be OR'd back in as appropriate. */ *nr = 0; global_ctrl = (*nr)++; @@ -5329,41 +5334,68 @@ static struct perf_guest_switch_msr *intel_guest_get_msrs(int *nr, void *data) return arr; } - if (!kvm_pmu || !x86_pmu.pebs_ept) + /* + * If the CPU doesn't support PEBS in the guest, then there's nothing + * more to do as disabling PMCs via PERF_GLOBAL_CTRL is sufficient on + * CPUs with guest/host isolation. + */ + if (!x86_pmu.pebs_ept) return arr; + /* + * Restrict guest PEBS events to counters that (a) perf supports, (b) + * the guest wants to use for PEBS, (c) are not excluded from counting + * in the guest, and (d) _are_ excluded from counting in the host. + */ + guest_pebs_mask = pebs_mask & intel_ctrl & kvm_pmu->pebs_enable & + ~cpuc->intel_ctrl_host_mask & + cpuc->intel_ctrl_guest_mask; + + /* + * Disable counters where the guest PMC is different than the host PMC + * being used on behalf of the guest, as the PEBS record includes + * PERF_GLOBAL_STATUS, i.e. the guest will see overflow status for the + * wrong counter(s). + */ + guest_pebs_mask &= ~kvm_pmu->host_cross_mapped_mask; + + /* + * FIXME: Allow guest and host usage of PEBS events to co-exist instead + * of disabling guest PEBS entirely if the host is using PEBS. + * What exactly goes wrong if guest and host are using PEBS is + * unknown. + */ + if (pebs_mask & ~cpuc->intel_ctrl_guest_mask) + guest_pebs_mask = 0; + + /* + * Context switch DS_AREA and PEBS_DATA_CFG if and only if PEBS will be + * active in the guest; if no records will be generated while the guest + * is running, then simply keep the host values resident in hardware. + */ arr[(*nr)++] = (struct perf_guest_switch_msr){ .msr = MSR_IA32_DS_AREA, .host = (unsigned long)cpuc->ds, - .guest = kvm_pmu->ds_area, + .guest = guest_pebs_mask ? kvm_pmu->ds_area : (unsigned long)cpuc->ds, }; if (x86_pmu.intel_cap.pebs_baseline) { arr[(*nr)++] = (struct perf_guest_switch_msr){ .msr = MSR_PEBS_DATA_CFG, .host = cpuc->active_pebs_data_cfg, - .guest = kvm_pmu->pebs_data_cfg, + .guest = guest_pebs_mask ? kvm_pmu->pebs_data_cfg : + cpuc->active_pebs_data_cfg, }; } - pebs_enable = (*nr)++; - arr[pebs_enable] = (struct perf_guest_switch_msr){ - .msr = MSR_IA32_PEBS_ENABLE, - .host = cpuc->pebs_enabled & ~cpuc->intel_ctrl_guest_mask, - .guest = pebs_mask & ~cpuc->intel_ctrl_host_mask & kvm_pmu->pebs_enable, - }; - - if (arr[pebs_enable].host) { - /* Disable guest PEBS if host PEBS is enabled. */ - arr[pebs_enable].guest = 0; - } else { - /* Disable guest PEBS thoroughly for cross-mapped PEBS counters. */ - arr[pebs_enable].guest &= ~kvm_pmu->host_cross_mapped_mask; - arr[global_ctrl].guest &= ~kvm_pmu->host_cross_mapped_mask; - /* Set hw GLOBAL_CTRL bits for PEBS counter when it runs for guest */ - arr[global_ctrl].guest |= arr[pebs_enable].guest; - } - + /* + * Do NOT mess with PEBS_ENABLED. As above, disabling counters via + * PERF_GLOBAL_CTRL is sufficient, and loading a stale PEBS_ENABLED, + * e.g. on VM-Exit, can put the system in a bad state. Simply enable + * counters in PERF_GLOBAL_CTRL, as perf load PEBS_ENABLED with the + * full value, i.e. perf *also* relies on PERF_GLOBAL_CTRL. + */ + arr[global_ctrl].guest |= guest_pebs_mask; return arr; } @@ -8755,8 +8787,6 @@ __init int intel_pmu_init(void) /* Initialize Atom core specific PerfMon capabilities.*/ pmu = &x86_pmu.hybrid_pmu[X86_HYBRID_PMU_ATOM_IDX]; intel_pmu_init_arw(&pmu->pmu); - - intel_pmu_pebs_data_source_lnl(); break; default: diff --git a/arch/x86/events/intel/ds.c b/arch/x86/events/intel/ds.c index 3dcd7dfc92b1..ccd98c463558 100644 --- a/arch/x86/events/intel/ds.c +++ b/arch/x86/events/intel/ds.c @@ -277,8 +277,8 @@ static u64 pnc_pebs_l2_hit_data_source[PNC_PEBS_DATA_SOURCE_MAX] = { 0, /* 0x06: Reserved */ OP_LH | P(LVL, L2) | LEVEL(L2) | P(SNOOP, HIT), /* 0x07: L2 Hit Snoop HIT */ OP_LH | P(LVL, L2) | LEVEL(L2) | P(SNOOP, HITM), /* 0x08: L2 Hit Snoop Hit Modified */ - OP_LH | P(LVL, L2) | LEVEL(L2) | P(SNOOP, MISS), /* 0x09: Prefetch Promotion */ - OP_LH | P(LVL, L2) | LEVEL(L2) | P(SNOOP, MISS), /* 0x0a: Cross Core Prefetch Promotion */ + OP_LH | P(LVL, L2) | LEVEL(L2) | P(SNOOP, NONE), /* 0x09: Prefetch Promotion */ + OP_LH | P(LVL, L2) | LEVEL(L2) | P(SNOOP, NONE), /* 0x0a: Cross Core Prefetch Promotion */ 0, /* 0x0b: Reserved */ 0, /* 0x0c: Reserved */ 0, /* 0x0d: Reserved */ @@ -455,6 +455,7 @@ static inline void pebs_set_tlb_lock(u64 *val, bool tlb, bool lock) static u64 __grt_latency_data(struct perf_event *event, u64 status, u8 dse, bool tlb, bool lock, bool blk) { + union perf_mem_data_src src; u64 val; WARN_ON_ONCE(is_hybrid() && @@ -470,7 +471,16 @@ static u64 __grt_latency_data(struct perf_event *event, u64 status, else val |= P(BLK, NA); - return val; + src.val = val; + + if (event->hw.flags & + (PERF_X86_EVENT_PEBS_LDLAT | PERF_X86_EVENT_PEBS_LD_HSW)) + src.mem_op = P(OP, LOAD); + if (event->hw.flags & + (PERF_X86_EVENT_PEBS_STLAT | PERF_X86_EVENT_PEBS_ST_HSW)) + src.mem_op = P(OP, STORE); + + return src.val; } u64 grt_latency_data(struct perf_event *event, u64 status) @@ -563,7 +573,11 @@ static u64 lnc_latency_data(struct perf_event *event, u64 status) val |= P(BLK, NA); src.val = val; - if (event->hw.flags & PERF_X86_EVENT_PEBS_ST_HSW) + if (event->hw.flags & + (PERF_X86_EVENT_PEBS_LDLAT | PERF_X86_EVENT_PEBS_LD_HSW)) + src.mem_op = P(OP, LOAD); + if (event->hw.flags & + (PERF_X86_EVENT_PEBS_STLAT | PERF_X86_EVENT_PEBS_ST_HSW)) src.mem_op = P(OP, STORE); return src.val; @@ -621,7 +635,11 @@ u64 pnc_latency_data(struct perf_event *event, u64 status) val |= P(BLK, NA); src.val = val; - if (event->hw.flags & PERF_X86_EVENT_PEBS_ST_HSW) + if (event->hw.flags & + (PERF_X86_EVENT_PEBS_LDLAT | PERF_X86_EVENT_PEBS_LD_HSW)) + src.mem_op = P(OP, LOAD); + if (event->hw.flags & + (PERF_X86_EVENT_PEBS_STLAT | PERF_X86_EVENT_PEBS_ST_HSW)) src.mem_op = P(OP, STORE); return src.val; @@ -1291,22 +1309,22 @@ struct event_constraint intel_glm_pebs_event_constraints[] = { struct event_constraint intel_grt_pebs_event_constraints[] = { /* Allow all events as PEBS with no flags */ - INTEL_HYBRID_LAT_CONSTRAINT(0x5d0, 0x3), - INTEL_HYBRID_LAT_CONSTRAINT(0x6d0, 0x3f), + INTEL_HYBRID_LDLAT_CONSTRAINT(0x5d0, 0x3), + INTEL_HYBRID_STLAT_CONSTRAINT(0x6d0, 0x3f), EVENT_CONSTRAINT_END }; struct event_constraint intel_cmt_pebs_event_constraints[] = { /* Allow all events as PEBS with no flags */ - INTEL_HYBRID_LAT_CONSTRAINT(0x5d0, 0x3), - INTEL_HYBRID_LAT_CONSTRAINT(0x6d0, 0xff), + INTEL_HYBRID_LDLAT_CONSTRAINT(0x5d0, 0x3), + INTEL_HYBRID_STLAT_CONSTRAINT(0x6d0, 0xff), EVENT_CONSTRAINT_END }; struct event_constraint intel_dkt_pebs_event_constraints[] = { /* Allow all events as PEBS with no flags */ - INTEL_HYBRID_LAT_CONSTRAINT(0x5d0, 0xff), - INTEL_HYBRID_LAT_CONSTRAINT(0x6d0, 0xff), + INTEL_HYBRID_LDLAT_CONSTRAINT(0x5d0, 0xff), + INTEL_HYBRID_STLAT_CONSTRAINT(0x6d0, 0xff), EVENT_CONSTRAINT_END }; @@ -1506,24 +1524,8 @@ struct event_constraint intel_lnc_pebs_event_constraints[] = { INTEL_FLAGS_UEVENT_CONSTRAINT(0x012a, 0x1), /* OCR.* events */ INTEL_FLAGS_UEVENT_CONSTRAINT(0x012b, 0x1), /* OCR.* events */ - INTEL_FLAGS_UEVENT_CONSTRAINT(0x04a4, 0x1), /* TOPDOWN.BAD_SPEC_SLOTS */ - INTEL_FLAGS_UEVENT_CONSTRAINT(0x08a4, 0x1), /* TOPDOWN.BR_MISPREDICT_SLOTS */ - INTEL_FLAGS_UEVENT_CONSTRAINT(0x10a4, 0x8), /* TOPDOWN.MEMORY_BOUND_SLOTS */ - INTEL_HYBRID_LDLAT_CONSTRAINT(0x1cd, 0x3fc), INTEL_HYBRID_STLAT_CONSTRAINT(0x2cd, 0x3), - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x11d0, 0xf), /* MEM_INST_RETIRED.STLB_MISS_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_ST(0x12d0, 0xf), /* MEM_INST_RETIRED.STLB_MISS_STORES */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x21d0, 0xf), /* MEM_INST_RETIRED.LOCK_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x41d0, 0xf), /* MEM_INST_RETIRED.SPLIT_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_ST(0x42d0, 0xf), /* MEM_INST_RETIRED.SPLIT_STORES */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x81d0, 0xf), /* MEM_INST_RETIRED.ALL_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_ST(0x82d0, 0xf), /* MEM_INST_RETIRED.ALL_STORES */ - INTEL_FLAGS_UEVENT_CONSTRAINT(0x87d0, 0x3ff), /* MEM_INST_RETIRED.ANY */ - - INTEL_FLAGS_EVENT_CONSTRAINT_DATALA_LD_RANGE(0xd1, 0xd4, 0xf), - - INTEL_FLAGS_EVENT_CONSTRAINT(0xd0, 0xf), /* * Everything else is handled by PMU_FL_PEBS_ALL, because we @@ -1539,18 +1541,6 @@ struct event_constraint intel_pnc_pebs_event_constraints[] = { INTEL_HYBRID_LDLAT_CONSTRAINT(0x1cd, 0xfc), INTEL_HYBRID_STLAT_CONSTRAINT(0x2cd, 0x3), - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x11d0, 0xf), /* MEM_INST_RETIRED.STLB_MISS_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_ST(0x12d0, 0xf), /* MEM_INST_RETIRED.STLB_MISS_STORES */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x21d0, 0xf), /* MEM_INST_RETIRED.LOCK_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x41d0, 0xf), /* MEM_INST_RETIRED.SPLIT_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_ST(0x42d0, 0xf), /* MEM_INST_RETIRED.SPLIT_STORES */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_LD(0x81d0, 0xf), /* MEM_INST_RETIRED.ALL_LOADS */ - INTEL_FLAGS_UEVENT_CONSTRAINT_DATALA_ST(0x82d0, 0xf), /* MEM_INST_RETIRED.ALL_STORES */ - - INTEL_FLAGS_EVENT_CONSTRAINT_DATALA_LD_RANGE(0xd1, 0xd4, 0xf), - - INTEL_FLAGS_EVENT_CONSTRAINT(0xd0, 0xf), - INTEL_FLAGS_EVENT_CONSTRAINT(0xd6, 0xf), /* * Everything else is handled by PMU_FL_PEBS_ALL, because we diff --git a/arch/x86/events/intel/lbr.c b/arch/x86/events/intel/lbr.c index 52d14927a191..1ed29d58f119 100644 --- a/arch/x86/events/intel/lbr.c +++ b/arch/x86/events/intel/lbr.c @@ -756,10 +756,10 @@ void intel_pmu_lbr_read_32(struct cpu_hw_events *cpuc) rdmsrq(x86_pmu.lbr_from + lbr_idx, msr_lastbranch.lbr); - perf_clear_branch_entry_bitfields(br); - - br->from = msr_lastbranch.from; - br->to = msr_lastbranch.to; + *br = (struct perf_branch_entry){ + .from = msr_lastbranch.from, + .to = msr_lastbranch.to, + }; br++; } cpuc->lbr_stack.nr = i; @@ -847,14 +847,15 @@ void intel_pmu_lbr_read_64(struct cpu_hw_events *cpuc) if (abort && x86_pmu.lbr_double_abort && out > 0) out--; - perf_clear_branch_entry_bitfields(br+out); - br[out].from = from; - br[out].to = to; - br[out].mispred = mis; - br[out].predicted = pred; - br[out].in_tx = in_tx; - br[out].abort = abort; - br[out].cycles = cycles; + br[out] = (struct perf_branch_entry){ + .from = from, + .to = to, + .mispred = mis, + .predicted = pred, + .in_tx = in_tx, + .abort = abort, + .cycles = cycles, + }; out++; } cpuc->lbr_stack.nr = out; @@ -905,6 +906,7 @@ static void intel_pmu_store_lbr(struct cpu_hw_events *cpuc, struct perf_branch_entry *e; struct lbr_entry *lbr; u64 from, to, info; + bool mispred; int i; for (i = 0; i < x86_pmu.lbr_nr; i++) { @@ -921,24 +923,27 @@ static void intel_pmu_store_lbr(struct cpu_hw_events *cpuc, to = rdlbr_to(i, lbr); info = rdlbr_info(i, lbr); - perf_clear_branch_entry_bitfields(e); - - e->from = from; - e->to = to; - e->mispred = get_lbr_mispred(info); - e->predicted = !e->mispred; - e->in_tx = !!(info & LBR_INFO_IN_TX); - e->abort = !!(info & LBR_INFO_ABORT); - e->cycles = get_lbr_cycles(info); - e->type = get_lbr_br_type(info); - - /* - * Leverage the reserved field of cpuc->lbr_entries[i] to - * temporarily store the branch counters information. - * The later code will decide what content can be disclosed - * to the perf tool. Pleae see intel_pmu_lbr_counters_reorder(). - */ - e->reserved = (info >> LBR_INFO_BR_CNTR_OFFSET) & LBR_INFO_BR_CNTR_FULL_MASK; + mispred = get_lbr_mispred(info); + + *e = (struct perf_branch_entry){ + .from = from, + .to = to, + .mispred = mispred, + .predicted = !mispred, + .in_tx = !!(info & LBR_INFO_IN_TX), + .abort = !!(info & LBR_INFO_ABORT), + .cycles = get_lbr_cycles(info), + .type = get_lbr_br_type(info), + /* + * Leverage the reserved field of + * cpuc->lbr_entries[i] to temporarily store the + * branch counters information. The later code will + * decide what content can be disclosed to the perf + * tool. Pleae see intel_pmu_lbr_counters_reorder(). + */ + .reserved = (info >> LBR_INFO_BR_CNTR_OFFSET) & + LBR_INFO_BR_CNTR_FULL_MASK, + }; } cpuc->lbr_stack.nr = i; diff --git a/arch/x86/events/perf_event.h b/arch/x86/events/perf_event.h index 680220d311a7..b577d6048bbb 100644 --- a/arch/x86/events/perf_event.h +++ b/arch/x86/events/perf_event.h @@ -517,10 +517,6 @@ struct cpu_hw_events { __EVENT_CONSTRAINT(c, n, INTEL_ARCH_EVENT_MASK|X86_ALL_EVENT_FLAGS, \ HWEIGHT(n), 0, PERF_X86_EVENT_PEBS_ST) -#define INTEL_HYBRID_LAT_CONSTRAINT(c, n) \ - __EVENT_CONSTRAINT(c, n, INTEL_ARCH_EVENT_MASK|X86_ALL_EVENT_FLAGS, \ - HWEIGHT(n), 0, PERF_X86_EVENT_PEBS_LAT_HYBRID) - #define INTEL_HYBRID_LDLAT_CONSTRAINT(c, n) \ __EVENT_CONSTRAINT(c, n, INTEL_ARCH_EVENT_MASK|X86_ALL_EVENT_FLAGS, \ HWEIGHT(n), 0, PERF_X86_EVENT_PEBS_LAT_HYBRID|PERF_X86_EVENT_PEBS_LD_HSW) diff --git a/arch/x86/include/asm/kvm-x86-nested-ops.h b/arch/x86/include/asm/kvm-x86-nested-ops.h new file mode 100644 index 000000000000..4b1be5bcecaa --- /dev/null +++ b/arch/x86/include/asm/kvm-x86-nested-ops.h @@ -0,0 +1,36 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +#if !defined(KVM_X86_NESTED_OP) || \ + !defined(KVM_X86_NESTED_OP_OPTIONAL) || \ + !defined(KVM_X86_NESTED_OP_OPTIONAL_RET0) +#error Missing one or more KVM_X86_NESTED_OP #defines +#else +/* + * KVM_X86_NESTED_OP() and KVM_X86_NESTED_OP_OPTIONAL() are used to help + * generate both DECLARE/DEFINE_STATIC_CALL() invocations and + * "static_call_update()" calls. + * + * KVM_X86_NESTED_OP_OPTIONAL() can be used for those functions that can have + * a NULL definition. KVM_X86_NESTED_OP_OPTIONAL_RET0() can be used likewise + * to make a definition optional, but in this case the default will + * be __static_call_return0. + */ +KVM_X86_NESTED_OP(leave_nested) +KVM_X86_NESTED_OP(is_exception_vmexit) +KVM_X86_NESTED_OP(check_events) +KVM_X86_NESTED_OP_OPTIONAL_RET0(has_events) +KVM_X86_NESTED_OP(triple_fault) +KVM_X86_NESTED_OP(get_state) +KVM_X86_NESTED_OP(set_state) +KVM_X86_NESTED_OP(get_nested_state_pages) +KVM_X86_NESTED_OP_OPTIONAL_RET0(write_log_dirty) +KVM_X86_NESTED_OP(translate_nested_gpa) +#ifdef CONFIG_KVM_HYPERV +KVM_X86_NESTED_OP_OPTIONAL(enable_evmcs) +KVM_X86_NESTED_OP_OPTIONAL_RET0(get_evmcs_version) +KVM_X86_NESTED_OP(hv_inject_synthetic_vmexit_post_tlb_flush) +#endif +#endif + +#undef KVM_X86_NESTED_OP +#undef KVM_X86_NESTED_OP_OPTIONAL +#undef KVM_X86_NESTED_OP_OPTIONAL_RET0 diff --git a/arch/x86/include/asm/kvm_host.h b/arch/x86/include/asm/kvm_host.h index 2dfd6199a10c..d02c6c42d2a0 100644 --- a/arch/x86/include/asm/kvm_host.h +++ b/arch/x86/include/asm/kvm_host.h @@ -2019,6 +2019,8 @@ struct kvm_x86_ops { }; struct kvm_x86_nested_ops { + bool enabled; + void (*leave_nested)(struct kvm_vcpu *vcpu); bool (*is_exception_vmexit)(struct kvm_vcpu *vcpu, u8 vector, u32 error_code); @@ -2075,6 +2077,14 @@ extern struct kvm_x86_ops kvm_x86_ops; #define KVM_X86_OP_OPTIONAL_RET0 KVM_X86_OP #include +#define kvm_nested_call(func) static_call(kvm_x86_nested_##func) + +#define KVM_X86_NESTED_OP(func) \ + DECLARE_STATIC_CALL(kvm_x86_nested_##func, *(((struct kvm_x86_nested_ops *)0)->func)); +#define KVM_X86_NESTED_OP_OPTIONAL KVM_X86_NESTED_OP +#define KVM_X86_NESTED_OP_OPTIONAL_RET0 KVM_X86_NESTED_OP +#include + int kvm_x86_vendor_init(struct kvm_x86_init_ops *ops); void kvm_x86_vendor_exit(void); @@ -2433,13 +2443,6 @@ static inline void kvm_inject_gp(struct kvm_vcpu *vcpu, u32 error_code) kvm_queue_exception_e(vcpu, GP_VECTOR, error_code); } -#define TSS_IOPB_BASE_OFFSET 0x66 -#define TSS_BASE_SIZE 0x68 -#define TSS_IOPB_SIZE (65536 / 8) -#define TSS_REDIRECTION_SIZE (256 / 8) -#define RMODE_TSS_SIZE \ - (TSS_BASE_SIZE + TSS_REDIRECTION_SIZE + TSS_IOPB_SIZE + 1) - enum { TASK_SWITCH_CALL = 0, TASK_SWITCH_IRET = 1, @@ -2462,12 +2465,7 @@ enum { # define kvm_memslots_for_spte_role(kvm, role) __kvm_memslots(kvm, 0) #endif -int kvm_cpu_has_injectable_intr(struct kvm_vcpu *v); -int kvm_cpu_has_interrupt(struct kvm_vcpu *vcpu); -int kvm_cpu_has_extint(struct kvm_vcpu *v); int kvm_arch_interrupt_allowed(struct kvm_vcpu *vcpu); -int kvm_cpu_get_extint(struct kvm_vcpu *v); -int kvm_cpu_get_interrupt(struct kvm_vcpu *v); void kvm_vcpu_reset(struct kvm_vcpu *vcpu, bool init_event); int kvm_pv_send_ipi(struct kvm *kvm, unsigned long ipi_bitmap_low, diff --git a/arch/x86/kernel/cpu/mce/core.c b/arch/x86/kernel/cpu/mce/core.c index cfb74be19994..1aebc8394fca 100644 --- a/arch/x86/kernel/cpu/mce/core.c +++ b/arch/x86/kernel/cpu/mce/core.c @@ -2108,6 +2108,9 @@ bool filter_mce(struct mce *m) static __always_inline void exc_machine_check_kernel(struct pt_regs *regs) { irqentry_state_t irq_state; + unsigned long dr7; + + dr7 = local_db_save(); WARN_ON_ONCE(user_mode(regs)); @@ -2116,20 +2119,26 @@ static __always_inline void exc_machine_check_kernel(struct pt_regs *regs) * mce_check_crashing_cpu() for details. */ if (mca_cfg.initialized && mce_check_crashing_cpu()) - return; + goto out; irq_state = irqentry_nmi_enter(regs); do_machine_check(regs); irqentry_nmi_exit(regs, irq_state); +out: + local_db_restore(dr7); } static __always_inline void exc_machine_check_user(struct pt_regs *regs) { + unsigned long dr7; + irqentry_enter_from_user_mode(regs); + dr7 = local_db_save(); do_machine_check(regs); + local_db_restore(dr7); irqentry_exit_to_user_mode(regs); } @@ -2138,21 +2147,13 @@ static __always_inline void exc_machine_check_user(struct pt_regs *regs) /* MCE hit kernel mode */ DEFINE_IDTENTRY_MCE(exc_machine_check) { - unsigned long dr7; - - dr7 = local_db_save(); exc_machine_check_kernel(regs); - local_db_restore(dr7); } /* The user mode variant. */ DEFINE_IDTENTRY_MCE_USER(exc_machine_check) { - unsigned long dr7; - - dr7 = local_db_save(); exc_machine_check_user(regs); - local_db_restore(dr7); } #ifdef CONFIG_X86_FRED @@ -2169,28 +2170,20 @@ DEFINE_IDTENTRY_MCE_USER(exc_machine_check) */ DEFINE_FREDENTRY_MCE(exc_machine_check) { - unsigned long dr7; - - dr7 = local_db_save(); if (user_mode(regs)) exc_machine_check_user(regs); else exc_machine_check_kernel(regs); - local_db_restore(dr7); } #endif #else /* 32bit unified entry point */ DEFINE_IDTENTRY_RAW(exc_machine_check) { - unsigned long dr7; - - dr7 = local_db_save(); if (user_mode(regs)) exc_machine_check_user(regs); else exc_machine_check_kernel(regs); - local_db_restore(dr7); } #endif diff --git a/arch/x86/kvm/hyperv.c b/arch/x86/kvm/hyperv.c index 5f43ae345f87..ddbd35ed255b 100644 --- a/arch/x86/kvm/hyperv.c +++ b/arch/x86/kvm/hyperv.c @@ -2419,7 +2419,7 @@ static int kvm_hv_hypercall_complete(struct kvm_vcpu *vcpu, u64 result) ret = kvm_skip_emulated_instruction(vcpu); if (tlb_lock_count) - kvm_x86_ops.nested_ops->hv_inject_synthetic_vmexit_post_tlb_flush(vcpu); + kvm_nested_call(hv_inject_synthetic_vmexit_post_tlb_flush)(vcpu); return ret; } @@ -2800,8 +2800,8 @@ int kvm_get_hv_cpuid(struct kvm_vcpu *vcpu, struct kvm_cpuid2 *cpuid, }; int i, nent = ARRAY_SIZE(cpuid_entries); - if (kvm_x86_ops.nested_ops->get_evmcs_version) - evmcs_ver = kvm_x86_ops.nested_ops->get_evmcs_version(vcpu); + if (kvm_x86_ops.nested_ops->enabled) + evmcs_ver = kvm_nested_call(get_evmcs_version)(vcpu); if (cpuid->nent < nent) return -E2BIG; diff --git a/arch/x86/kvm/irq.h b/arch/x86/kvm/irq.h index 34f4a78a7a01..1a84ea31e7fd 100644 --- a/arch/x86/kvm/irq.h +++ b/arch/x86/kvm/irq.h @@ -112,6 +112,12 @@ static inline int irqchip_in_kernel(struct kvm *kvm) return mode != KVM_IRQCHIP_NONE; } +int kvm_cpu_has_injectable_intr(struct kvm_vcpu *v); +int kvm_cpu_has_interrupt(struct kvm_vcpu *vcpu); +int kvm_cpu_has_extint(struct kvm_vcpu *v); +int kvm_cpu_get_extint(struct kvm_vcpu *v); +int kvm_cpu_get_interrupt(struct kvm_vcpu *v); + void kvm_inject_pending_timer_irqs(struct kvm_vcpu *vcpu); void kvm_inject_apic_timer_irqs(struct kvm_vcpu *vcpu); void kvm_apic_nmi_wd_deliver(struct kvm_vcpu *vcpu); diff --git a/arch/x86/kvm/mmu.h b/arch/x86/kvm/mmu.h index e1bb663ebbd5..bb6e01892852 100644 --- a/arch/x86/kvm/mmu.h +++ b/arch/x86/kvm/mmu.h @@ -308,9 +308,8 @@ static inline gpa_t kvm_translate_gpa(struct kvm_vcpu *vcpu, { if (mmu != &vcpu->arch.nested_mmu) return gpa; - return kvm_x86_ops.nested_ops->translate_nested_gpa(vcpu, gpa, access, - exception, - pte_access); + return kvm_nested_call(translate_nested_gpa)(vcpu, gpa, access, + exception, pte_access); } static inline bool kvm_has_mirrored_tdp(const struct kvm *kvm) diff --git a/arch/x86/kvm/mmu/mmu.c b/arch/x86/kvm/mmu/mmu.c index 51ba46aa6b7e..49a5fb103194 100644 --- a/arch/x86/kvm/mmu/mmu.c +++ b/arch/x86/kvm/mmu/mmu.c @@ -4600,37 +4600,6 @@ static bool kvm_arch_setup_async_pf(struct kvm_vcpu *vcpu, kvm_vcpu_gfn_to_hva(vcpu, fault->gfn), &arch); } -void kvm_arch_async_page_ready(struct kvm_vcpu *vcpu, struct kvm_async_pf *work) -{ - int r; - - if (WARN_ON_ONCE(work->arch.error_code & PFERR_PRIVATE_ACCESS)) - return; - - if ((vcpu->arch.mmu->root_role.direct != work->arch.direct_map) || - work->wakeup_all) - return; - - r = kvm_mmu_reload(vcpu); - if (unlikely(r)) - return; - - if (!vcpu->arch.mmu->root_role.direct && - work->arch.cr3 != kvm_mmu_get_guest_pgd(vcpu, vcpu->arch.mmu)) - return; - - r = kvm_mmu_do_page_fault(vcpu, work->cr2_or_gpa, work->arch.error_code, - true, NULL, NULL); - - /* - * Account fixed page faults, otherwise they'll never be counted, but - * ignore stats for all other return times. Page-ready "faults" aren't - * truly spurious and never trigger emulation - */ - if (r == RET_PF_FIXED) - vcpu->stat.pf_fixed++; -} - static void kvm_mmu_finish_page_fault(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault, int r) { @@ -4988,7 +4957,7 @@ static int kvm_tdp_mmu_page_fault(struct kvm_vcpu *vcpu, } #endif -int kvm_tdp_page_fault(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault) +static int kvm_tdp_page_fault(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault) { #ifdef CONFIG_X86_64 if (tdp_mmu_enabled) @@ -4998,6 +4967,71 @@ int kvm_tdp_page_fault(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault) return direct_page_fault(vcpu, fault); } +static int kvm_mmu_do_page_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, + u64 err, bool prefetch, int *emulation_type, + u8 *level) +{ + struct kvm_page_fault fault = { + .addr = cr2_or_gpa, + .error_code = err, + .exec = err & PFERR_FETCH_MASK, + .write = err & PFERR_WRITE_MASK, + .present = err & PFERR_PRESENT_MASK, + .rsvd = err & PFERR_RSVD_MASK, + .user = err & PFERR_USER_MASK, + .prefetch = prefetch, + .is_tdp = likely(vcpu->arch.mmu->page_fault == kvm_tdp_page_fault), + .nx_huge_page_workaround_enabled = + is_nx_huge_page_enabled(vcpu->kvm), + + .max_level = KVM_MAX_HUGEPAGE_LEVEL, + .req_level = PG_LEVEL_4K, + .goal_level = PG_LEVEL_4K, + .is_private = err & PFERR_PRIVATE_ACCESS, + + .pfn = KVM_PFN_ERR_FAULT, + }; + int r; + + if (vcpu->arch.mmu->root_role.direct) { + /* + * Things like memslots don't understand the concept of a shared + * bit. Strip it so that the GFN can be used like normal, and the + * fault.addr can be used when the shared bit is needed. + */ + fault.gfn = gpa_to_gfn(fault.addr) & ~kvm_gfn_direct_bits(vcpu->kvm); + fault.slot = kvm_vcpu_gfn_to_memslot(vcpu, fault.gfn); + } + + /* + * With retpoline being active an indirect call is rather expensive, + * so do a direct call in the most common case. + */ + if (IS_ENABLED(CONFIG_MITIGATION_RETPOLINE) && fault.is_tdp) + r = kvm_tdp_page_fault(vcpu, &fault); + else + r = vcpu->arch.mmu->page_fault(vcpu, &fault); + + /* + * Not sure what's happening, but punt to userspace and hope that + * they can fix it by changing memory to shared, or they can + * provide a better error. + */ + if (r == RET_PF_EMULATE && fault.is_private) { + pr_warn_ratelimited("kvm: unexpected emulation request on private memory\n"); + kvm_mmu_prepare_memory_fault_exit(vcpu, &fault); + return -EFAULT; + } + + if (fault.write_fault_to_shadow_pgtable && emulation_type) + *emulation_type |= EMULTYPE_WRITE_PF_TO_SP; + if (level) + *level = fault.goal_level; + + return r; +} + + static int kvm_tdp_page_prefault(struct kvm_vcpu *vcpu, gpa_t gpa, u64 error_code, u8 *level) { @@ -5088,6 +5122,37 @@ long kvm_arch_vcpu_pre_fault_memory(struct kvm_vcpu *vcpu, return min(range->size, end - range->gpa); } +void kvm_arch_async_page_ready(struct kvm_vcpu *vcpu, struct kvm_async_pf *work) +{ + int r; + + if (WARN_ON_ONCE(work->arch.error_code & PFERR_PRIVATE_ACCESS)) + return; + + if ((vcpu->arch.mmu->root_role.direct != work->arch.direct_map) || + work->wakeup_all) + return; + + r = kvm_mmu_reload(vcpu); + if (unlikely(r)) + return; + + if (!vcpu->arch.mmu->root_role.direct && + work->arch.cr3 != kvm_mmu_get_guest_pgd(vcpu, vcpu->arch.mmu)) + return; + + r = kvm_mmu_do_page_fault(vcpu, work->cr2_or_gpa, work->arch.error_code, + true, NULL, NULL); + + /* + * Account fixed page faults, otherwise they'll never be counted, but + * ignore stats for all other return times. Page-ready "faults" aren't + * truly spurious and never trigger emulation + */ + if (r == RET_PF_FIXED) + vcpu->stat.pf_fixed++; +} + #ifdef CONFIG_KVM_GUEST_MEMFD static void kvm_assert_gmem_invalidate_lock_held(struct kvm_memory_slot *slot) { diff --git a/arch/x86/kvm/mmu/mmu_internal.h b/arch/x86/kvm/mmu/mmu_internal.h index 73cdcbccc89e..c29002c60126 100644 --- a/arch/x86/kvm/mmu/mmu_internal.h +++ b/arch/x86/kvm/mmu/mmu_internal.h @@ -290,8 +290,6 @@ struct kvm_page_fault { bool write_fault_to_shadow_pgtable; }; -int kvm_tdp_page_fault(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault); - /* * Return values of handle_mmio_page_fault(), mmu.page_fault(), fast_page_fault(), * and of course kvm_mmu_do_page_fault(). @@ -337,70 +335,6 @@ static inline void kvm_mmu_prepare_memory_fault_exit(struct kvm_vcpu *vcpu, fault->is_private); } -static inline int kvm_mmu_do_page_fault(struct kvm_vcpu *vcpu, gpa_t cr2_or_gpa, - u64 err, bool prefetch, - int *emulation_type, u8 *level) -{ - struct kvm_page_fault fault = { - .addr = cr2_or_gpa, - .error_code = err, - .exec = err & PFERR_FETCH_MASK, - .write = err & PFERR_WRITE_MASK, - .present = err & PFERR_PRESENT_MASK, - .rsvd = err & PFERR_RSVD_MASK, - .user = err & PFERR_USER_MASK, - .prefetch = prefetch, - .is_tdp = likely(vcpu->arch.mmu->page_fault == kvm_tdp_page_fault), - .nx_huge_page_workaround_enabled = - is_nx_huge_page_enabled(vcpu->kvm), - - .max_level = KVM_MAX_HUGEPAGE_LEVEL, - .req_level = PG_LEVEL_4K, - .goal_level = PG_LEVEL_4K, - .is_private = err & PFERR_PRIVATE_ACCESS, - - .pfn = KVM_PFN_ERR_FAULT, - }; - int r; - - if (vcpu->arch.mmu->root_role.direct) { - /* - * Things like memslots don't understand the concept of a shared - * bit. Strip it so that the GFN can be used like normal, and the - * fault.addr can be used when the shared bit is needed. - */ - fault.gfn = gpa_to_gfn(fault.addr) & ~kvm_gfn_direct_bits(vcpu->kvm); - fault.slot = kvm_vcpu_gfn_to_memslot(vcpu, fault.gfn); - } - - /* - * With retpoline being active an indirect call is rather expensive, - * so do a direct call in the most common case. - */ - if (IS_ENABLED(CONFIG_MITIGATION_RETPOLINE) && fault.is_tdp) - r = kvm_tdp_page_fault(vcpu, &fault); - else - r = vcpu->arch.mmu->page_fault(vcpu, &fault); - - /* - * Not sure what's happening, but punt to userspace and hope that - * they can fix it by changing memory to shared, or they can - * provide a better error. - */ - if (r == RET_PF_EMULATE && fault.is_private) { - pr_warn_ratelimited("kvm: unexpected emulation request on private memory\n"); - kvm_mmu_prepare_memory_fault_exit(vcpu, &fault); - return -EFAULT; - } - - if (fault.write_fault_to_shadow_pgtable && emulation_type) - *emulation_type |= EMULTYPE_WRITE_PF_TO_SP; - if (level) - *level = fault.goal_level; - - return r; -} - int kvm_mmu_max_mapping_level(struct kvm *kvm, struct kvm_page_fault *fault, const struct kvm_memory_slot *slot, gfn_t gfn); void kvm_mmu_hugepage_adjust(struct kvm_vcpu *vcpu, struct kvm_page_fault *fault); diff --git a/arch/x86/kvm/mmu/paging_tmpl.h b/arch/x86/kvm/mmu/paging_tmpl.h index 1ba840a73b7a..8d2c46bc6c0f 100644 --- a/arch/x86/kvm/mmu/paging_tmpl.h +++ b/arch/x86/kvm/mmu/paging_tmpl.h @@ -233,7 +233,7 @@ static int FNAME(update_accessed_dirty_bits)(struct kvm_vcpu *vcpu, !(pte & PT_GUEST_DIRTY_MASK)) { trace_kvm_mmu_set_dirty_bit(table_gfn, index, sizeof(pte)); #if PTTYPE == PTTYPE_EPT - if (kvm_x86_ops.nested_ops->write_log_dirty(vcpu, addr)) + if (kvm_nested_call(write_log_dirty)(vcpu, addr)) return -EINVAL; #endif pte |= PT_GUEST_DIRTY_MASK; diff --git a/arch/x86/kvm/pmu.c b/arch/x86/kvm/pmu.c index dd1c57593f48..7fedf25cf46f 100644 --- a/arch/x86/kvm/pmu.c +++ b/arch/x86/kvm/pmu.c @@ -812,14 +812,6 @@ void kvm_pmu_deliver_pmi(struct kvm_vcpu *vcpu) bool kvm_pmu_is_valid_msr(struct kvm_vcpu *vcpu, u32 msr) { - switch (msr) { - case MSR_CORE_PERF_GLOBAL_STATUS: - case MSR_CORE_PERF_GLOBAL_CTRL: - case MSR_CORE_PERF_GLOBAL_OVF_CTRL: - return kvm_pmu_has_perf_global_ctrl(vcpu_to_pmu(vcpu)); - default: - break; - } return kvm_pmu_call(msr_idx_to_pmc)(vcpu, msr) || kvm_pmu_call(is_valid_msr)(vcpu, msr); } diff --git a/arch/x86/kvm/svm/nested.c b/arch/x86/kvm/svm/nested.c index 3e6c671a8dc2..c1485c3e691c 100644 --- a/arch/x86/kvm/svm/nested.c +++ b/arch/x86/kvm/svm/nested.c @@ -23,6 +23,7 @@ #include "kvm_emulate.h" #include "trace.h" +#include "irq.h" #include "mmu.h" #include "x86.h" #include "smm.h" diff --git a/arch/x86/kvm/svm/sev.c b/arch/x86/kvm/svm/sev.c index 1ed78f770340..62561da83107 100644 --- a/arch/x86/kvm/svm/sev.c +++ b/arch/x86/kvm/svm/sev.c @@ -2048,6 +2048,12 @@ static void sev_migrate_from(struct kvm *dst_kvm, struct kvm *src_kvm) src->pages_locked = 0; src->es_active = false; + /* + * Do cache maintenance on the source VM as it is no longer an SEV VM, + * i.e. memory reclaim flows won't trigger cache maintenance on the VM. + */ + sev_writeback_caches(src_kvm); + list_cut_before(&dst->regions_list, &src->regions_list, &src->regions_list); mutex_lock(&sev_mirror_lock); @@ -2187,6 +2193,10 @@ int sev_vm_move_enc_context_from(struct kvm *kvm, unsigned int source_fd) * the set of CPUs from the source. If a CPU was used to run a vCPU in * the source VM but is never used for the destination VM, then the CPU * can only have cached memory that was accessible to the source VM. + * Furthermore, KVM *must* perform cache maintenance on the source VM, + * as the source VM may have access to memory that the destination VM + * does not, i.e. KVM could skip flushes if memory is reclaimed from + * the old VM but not the new VM. */ if (!zalloc_cpumask_var(&dst_sev->have_run_cpus, GFP_KERNEL_ACCOUNT)) { ret = -ENOMEM; @@ -2981,13 +2991,17 @@ void sev_vm_destroy(struct kvm *kvm) struct list_head *head = &sev->regions_list; struct list_head *pos, *q; + /* + * Free the mask even if the VM is not *currently* an SEV VM, as it may + * have been an SEV VM prior to intra-host migration. + */ + free_cpumask_var(sev->have_run_cpus); + if (!sev_guest(kvm)) return; WARN_ON(!list_empty(&sev->mirror_vms)); - free_cpumask_var(sev->have_run_cpus); - /* * If this is a mirror VM, remove it from the owner's list of a mirrors * and skip ASID cleanup (the ASID is tied to the lifetime of the owner). diff --git a/arch/x86/kvm/svm/svm.c b/arch/x86/kvm/svm/svm.c index 03fb9d7231c8..d1423bab1a3a 100644 --- a/arch/x86/kvm/svm/svm.c +++ b/arch/x86/kvm/svm/svm.c @@ -5662,6 +5662,7 @@ static __init int svm_hardware_setup(void) if (r) return r; } + svm_nested_ops.enabled = nested; /* * KVM's MMU doesn't support using 2-level paging for itself, and thus diff --git a/arch/x86/kvm/tss.h b/arch/x86/kvm/tss.h index 3f9150125e70..117bf8bec07d 100644 --- a/arch/x86/kvm/tss.h +++ b/arch/x86/kvm/tss.h @@ -57,4 +57,11 @@ struct tss_segment_16 { u16 ldt; }; +#define TSS_IOPB_BASE_OFFSET 0x66 +#define TSS_BASE_SIZE 0x68 +#define TSS_IOPB_SIZE (65536 / 8) +#define TSS_REDIRECTION_SIZE (256 / 8) +#define RMODE_TSS_SIZE \ + (TSS_BASE_SIZE + TSS_REDIRECTION_SIZE + TSS_IOPB_SIZE + 1) + #endif diff --git a/arch/x86/kvm/vmx/nested.c b/arch/x86/kvm/vmx/nested.c index ed28cf6b6efc..cec7eb646a1e 100644 --- a/arch/x86/kvm/vmx/nested.c +++ b/arch/x86/kvm/vmx/nested.c @@ -11,6 +11,7 @@ #include "x86.h" #include "cpuid.h" #include "hyperv.h" +#include "irq.h" #include "mmu.h" #include "nested.h" #include "pmu.h" diff --git a/arch/x86/kvm/vmx/pmu_intel.c b/arch/x86/kvm/vmx/pmu_intel.c index 1944939c4139..e65ba7e85611 100644 --- a/arch/x86/kvm/vmx/pmu_intel.c +++ b/arch/x86/kvm/vmx/pmu_intel.c @@ -187,6 +187,9 @@ static bool intel_is_valid_msr(struct kvm_vcpu *vcpu, u32 msr) int ret; switch (msr) { + case MSR_CORE_PERF_GLOBAL_STATUS: + case MSR_CORE_PERF_GLOBAL_CTRL: + case MSR_CORE_PERF_GLOBAL_OVF_CTRL: case MSR_CORE_PERF_FIXED_CTR_CTRL: return kvm_pmu_has_perf_global_ctrl(pmu); case MSR_IA32_PEBS_ENABLE: diff --git a/arch/x86/kvm/vmx/vmx.c b/arch/x86/kvm/vmx/vmx.c index b8d745f6fd22..61243d16febf 100644 --- a/arch/x86/kvm/vmx/vmx.c +++ b/arch/x86/kvm/vmx/vmx.c @@ -72,6 +72,7 @@ #include "x86.h" #include "x86_ops.h" #include "smm.h" +#include "tss.h" #include "vmx_onhyperv.h" #include "vmenter.h" #include "posted_intr.h" @@ -8770,6 +8771,7 @@ __init int vmx_hardware_setup(void) if (r) return r; } + vmx_nested_ops.enabled = nested; kvm_set_posted_intr_wakeup_handler(pi_wakeup_handler); diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 9822ff0450bb..61db05ff023c 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -148,6 +148,13 @@ EXPORT_STATIC_CALL_GPL(kvm_x86_get_cs_db_l_bits); EXPORT_STATIC_CALL_GPL(kvm_x86_cache_reg); EXPORT_STATIC_CALL_GPL(kvm_x86_get_cpl); +#define KVM_X86_NESTED_OP(func) \ + DEFINE_STATIC_CALL_NULL(kvm_x86_nested_##func, \ + *(((struct kvm_x86_nested_ops *)0)->func)); +#define KVM_X86_NESTED_OP_OPTIONAL KVM_X86_NESTED_OP +#define KVM_X86_NESTED_OP_OPTIONAL_RET0 KVM_X86_NESTED_OP +#include + static bool __read_mostly ignore_msrs = 0; module_param(ignore_msrs, bool, 0644); @@ -833,7 +840,7 @@ static void kvm_multiple_exception(struct kvm_vcpu *vcpu, unsigned int nr, * wants to intercept the exception. */ if (is_guest_mode(vcpu) && - kvm_x86_ops.nested_ops->is_exception_vmexit(vcpu, nr, error_code)) { + kvm_nested_call(is_exception_vmexit)(vcpu, nr, error_code)) { kvm_queue_exception_vmexit(vcpu, nr, has_error, error_code, has_payload, payload); return; @@ -4517,15 +4524,16 @@ int kvm_vm_ioctl_check_extension(struct kvm *kvm, long ext) r &= ~KVM_X2APIC_ENABLE_SUPPRESS_EOI_BROADCAST; break; case KVM_CAP_NESTED_STATE: - r = kvm_x86_ops.nested_ops->get_state ? - kvm_x86_ops.nested_ops->get_state(NULL, NULL, 0) : 0; + r = kvm_x86_ops.nested_ops->enabled ? + kvm_nested_call(get_state)(NULL, NULL, 0) : 0; break; #ifdef CONFIG_KVM_HYPERV case KVM_CAP_HYPERV_DIRECT_TLBFLUSH: r = kvm_x86_ops.enable_l2_tlb_flush != NULL; break; case KVM_CAP_HYPERV_ENLIGHTENED_VMCS: - r = kvm_x86_ops.nested_ops->enable_evmcs != NULL; + r = kvm_x86_ops.nested_ops->enabled && + kvm_x86_ops.nested_ops->enable_evmcs != NULL; break; #endif case KVM_CAP_SMALLER_MAXPHYADDR: @@ -5567,9 +5575,10 @@ static int kvm_vcpu_ioctl_enable_cap(struct kvm_vcpu *vcpu, uint16_t vmcs_version; void __user *user_ptr; - if (!kvm_x86_ops.nested_ops->enable_evmcs) + if (!kvm_x86_ops.nested_ops->enabled || + !kvm_x86_ops.nested_ops->enable_evmcs) return -ENOTTY; - r = kvm_x86_ops.nested_ops->enable_evmcs(vcpu, &vmcs_version); + r = kvm_nested_call(enable_evmcs)(vcpu, &vmcs_version); if (!r) { user_ptr = (void __user *)(uintptr_t)cap->args[0]; if (copy_to_user(user_ptr, &vmcs_version, @@ -6067,7 +6076,7 @@ long kvm_arch_vcpu_ioctl(struct file *filp, u32 user_data_size; r = -EINVAL; - if (!kvm_x86_ops.nested_ops->get_state) + if (!kvm_x86_ops.nested_ops->enabled) break; BUILD_BUG_ON(sizeof(user_data_size) != sizeof(user_kvm_nested_state->size)); @@ -6075,8 +6084,7 @@ long kvm_arch_vcpu_ioctl(struct file *filp, if (get_user(user_data_size, &user_kvm_nested_state->size)) break; - r = kvm_x86_ops.nested_ops->get_state(vcpu, user_kvm_nested_state, - user_data_size); + r = kvm_nested_call(get_state)(vcpu, user_kvm_nested_state, user_data_size); if (r < 0) break; @@ -6097,7 +6105,7 @@ long kvm_arch_vcpu_ioctl(struct file *filp, int idx; r = -EINVAL; - if (!kvm_x86_ops.nested_ops->set_state) + if (!kvm_x86_ops.nested_ops->enabled) break; r = -EFAULT; @@ -6120,7 +6128,7 @@ long kvm_arch_vcpu_ioctl(struct file *filp, break; idx = srcu_read_lock(&vcpu->kvm->srcu); - r = kvm_x86_ops.nested_ops->set_state(vcpu, user_kvm_nested_state, &kvm_state); + r = kvm_nested_call(set_state)(vcpu, user_kvm_nested_state, &kvm_state); srcu_read_unlock(&vcpu->kvm->srcu, idx); break; } @@ -9558,6 +9566,8 @@ static void kvm_setup_efer_caps(void) static inline void kvm_ops_update(struct kvm_x86_init_ops *ops) { + const struct kvm_x86_nested_ops *nested_ops = ops->runtime_ops->nested_ops; + memcpy(&kvm_x86_ops, ops->runtime_ops, sizeof(kvm_x86_ops)); #define __KVM_X86_OP(func) \ @@ -9571,6 +9581,17 @@ static inline void kvm_ops_update(struct kvm_x86_init_ops *ops) #include #undef __KVM_X86_OP +#define __KVM_X86_NESTED_OP(func) \ + static_call_update(kvm_x86_nested_##func, nested_ops->func); +#define KVM_X86_NESTED_OP(func) \ + WARN_ON(!nested_ops->func); __KVM_X86_NESTED_OP(func) +#define KVM_X86_NESTED_OP_OPTIONAL __KVM_X86_NESTED_OP +#define KVM_X86_NESTED_OP_OPTIONAL_RET0(func) \ + static_call_update(kvm_x86_nested_##func, (void *)nested_ops->func ? : \ + (void *)__static_call_return0); +#include +#undef __KVM_X86_NESTED_OP + kvm_pmu_ops_update(ops->pmu_ops); } @@ -10107,11 +10128,11 @@ static void post_kvm_run_save(struct kvm_vcpu *vcpu) int kvm_check_nested_events(struct kvm_vcpu *vcpu) { if (kvm_test_request(KVM_REQ_TRIPLE_FAULT, vcpu)) { - kvm_x86_ops.nested_ops->triple_fault(vcpu); + kvm_nested_call(triple_fault)(vcpu); return 1; } - return kvm_x86_ops.nested_ops->check_events(vcpu); + return kvm_nested_call(check_events)(vcpu); } static void kvm_inject_exception(struct kvm_vcpu *vcpu) @@ -10349,9 +10370,7 @@ static int kvm_check_and_inject_events(struct kvm_vcpu *vcpu, kvm_x86_call(enable_irq_window)(vcpu); } - if (is_guest_mode(vcpu) && - kvm_x86_ops.nested_ops->has_events && - kvm_x86_ops.nested_ops->has_events(vcpu, true)) + if (is_guest_mode(vcpu) && kvm_nested_call(has_events)(vcpu, true)) *req_immediate_exit = true; /* @@ -10674,7 +10693,7 @@ static int vcpu_enter_guest(struct kvm_vcpu *vcpu) } if (kvm_check_request(KVM_REQ_GET_NESTED_STATE_PAGES, vcpu)) { - if (unlikely(!kvm_x86_ops.nested_ops->get_nested_state_pages(vcpu))) { + if (unlikely(!kvm_nested_call(get_nested_state_pages)(vcpu))) { r = 0; goto out; } @@ -10726,7 +10745,7 @@ static int vcpu_enter_guest(struct kvm_vcpu *vcpu) } if (kvm_test_request(KVM_REQ_TRIPLE_FAULT, vcpu)) { if (is_guest_mode(vcpu)) - kvm_x86_ops.nested_ops->triple_fault(vcpu); + kvm_nested_call(triple_fault)(vcpu); if (kvm_check_request(KVM_REQ_TRIPLE_FAULT, vcpu)) { vcpu->run->exit_reason = KVM_EXIT_SHUTDOWN; @@ -11146,9 +11165,7 @@ bool kvm_vcpu_has_events(struct kvm_vcpu *vcpu) if (kvm_hv_has_stimer_pending(vcpu)) return true; - if (is_guest_mode(vcpu) && - kvm_x86_ops.nested_ops->has_events && - kvm_x86_ops.nested_ops->has_events(vcpu, false)) + if (is_guest_mode(vcpu) && kvm_nested_call(has_events)(vcpu, false)) return true; if (kvm_xen_has_pending_events(vcpu)) @@ -11573,8 +11590,7 @@ int kvm_arch_vcpu_ioctl_run(struct kvm_vcpu *vcpu) * a pending VM-Exit if L1 wants to intercept the exception. */ if (vcpu->arch.exception_from_userspace && is_guest_mode(vcpu) && - kvm_x86_ops.nested_ops->is_exception_vmexit(vcpu, ex->vector, - ex->error_code)) { + kvm_nested_call(is_exception_vmexit)(vcpu, ex->vector, ex->error_code)) { kvm_queue_exception_vmexit(vcpu, ex->vector, ex->has_error_code, ex->error_code, ex->has_payload, ex->payload); diff --git a/arch/x86/kvm/x86.h b/arch/x86/kvm/x86.h index bd2699bd3fe8..9973f905955a 100644 --- a/arch/x86/kvm/x86.h +++ b/arch/x86/kvm/x86.h @@ -151,7 +151,7 @@ int kvm_check_nested_events(struct kvm_vcpu *vcpu); /* Forcibly leave the nested mode in cases like a vCPU reset */ static inline void kvm_leave_nested(struct kvm_vcpu *vcpu) { - kvm_x86_ops.nested_ops->leave_nested(vcpu); + kvm_nested_call(leave_nested)(vcpu); } /* diff --git a/arch/x86/pci/fixup.c b/arch/x86/pci/fixup.c index b301c6c8df75..795d81c7a4de 100644 --- a/arch/x86/pci/fixup.c +++ b/arch/x86/pci/fixup.c @@ -886,6 +886,105 @@ static void quirk_clear_strap_no_soft_reset_dev2_f0(struct pci_dev *dev) } } DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_AMD, 0x15b8, quirk_clear_strap_no_soft_reset_dev2_f0); + +/* + * Enhanced atomic operations can cause corruption with 64-bit DMA + * on these devices. + */ +#define RX_ENH_ATOMIC_EN BIT(8) + +static const u32 nbio_7_7_pcie_smn_addrs[] = { + 0x111401d0, + 0x111411d0, + 0x111421d0, + 0x111431d0, + 0x111441d0, + 0x112401d0, + 0x112411d0, + 0x112421d0, + 0x112431d0, + 0x112441d0, + 0x112451d0, + 0x113401d0, + 0x114401d0, +}; + +static const u32 nbio_7_11_pcie_smn_addrs[] = { + 0x112401d0, + 0x112411d0, + 0x112421d0, + 0x112431d0, + 0x112441d0, + 0x112451d0, + 0x113401d0, + 0x113411d0, + 0x113421d0, + 0x113431d0, + 0x113441d0, + 0x113451d0, +}; + +static void quirk_amd_nbio_enhanced_atomic(struct pci_dev *host_bridge, + const u32 *smn_addrs, + size_t nr_smn_addrs) +{ + bool changed = false; + size_t i; + u32 data; + int ret; + + for (i = 0; i < nr_smn_addrs; i++) { + ret = amd_smn_read(0, smn_addrs[i], &data); + if (ret) + continue; + if (!(data & RX_ENH_ATOMIC_EN)) + continue; + data = data & ~RX_ENH_ATOMIC_EN; + ret = amd_smn_write(0, smn_addrs[i], data); + if (ret) + continue; + if (changed) + continue; + ret = amd_smn_read(0, smn_addrs[i], &data); + if (ret) + continue; + if (data & RX_ENH_ATOMIC_EN) + continue; + changed = true; + } + + if (changed) + pci_info(host_bridge, "enhanced atomics disabled\n"); +} + +static void quirk_amd_nbio_7_7_disable_enhanced_atomic(struct pci_dev *dev) +{ + quirk_amd_nbio_enhanced_atomic(dev, nbio_7_7_pcie_smn_addrs, + ARRAY_SIZE(nbio_7_7_pcie_smn_addrs)); +} + +static void quirk_amd_nbio_7_11_disable_enhanced_atomic(struct pci_dev *dev) +{ + quirk_amd_nbio_enhanced_atomic(dev, nbio_7_11_pcie_smn_addrs, + ARRAY_SIZE(nbio_7_11_pcie_smn_addrs)); +} + +/* Phoenix, Hawk Point (NBIO 7.7) */ +DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_AMD, 0x14E8, + quirk_amd_nbio_7_7_disable_enhanced_atomic); +DECLARE_PCI_FIXUP_RESUME(PCI_VENDOR_ID_AMD, 0x14E8, + quirk_amd_nbio_7_7_disable_enhanced_atomic); + +/* Strix, Krackan, Strix Halo (NBIO 7.11) */ +DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_AMD, 0x1507, + quirk_amd_nbio_7_11_disable_enhanced_atomic); +DECLARE_PCI_FIXUP_RESUME(PCI_VENDOR_ID_AMD, 0x1507, + quirk_amd_nbio_7_11_disable_enhanced_atomic); +DECLARE_PCI_FIXUP_FINAL(PCI_VENDOR_ID_AMD, 0x1122, + quirk_amd_nbio_7_11_disable_enhanced_atomic); +DECLARE_PCI_FIXUP_RESUME(PCI_VENDOR_ID_AMD, 0x1122, + quirk_amd_nbio_7_11_disable_enhanced_atomic); + #endif /* diff --git a/block/blk-zoned.c b/block/blk-zoned.c index ca30caec838e..dd6bc679210c 100644 --- a/block/blk-zoned.c +++ b/block/blk-zoned.c @@ -2048,12 +2048,17 @@ static int disk_revalidate_zone_resources(struct gendisk *disk, struct blk_revalidate_zone_args *args) { struct queue_limits *lim = &disk->queue->limits; + unsigned long long nr_zones; unsigned int pool_size; int ret = 0; args->disk = disk; - args->nr_zones = - DIV_ROUND_UP_ULL(get_capacity(disk), lim->chunk_sectors); + nr_zones = DIV_ROUND_UP_ULL(get_capacity(disk), lim->chunk_sectors); + if (nr_zones > UINT_MAX) { + pr_warn("%s: Too many zones (%llu)\n", disk->disk_name, nr_zones); + return -EINVAL; + } + args->nr_zones = nr_zones; /* Cached zone conditions: 1 byte per zone */ args->zones_cond = kzalloc(args->nr_zones, GFP_NOIO); @@ -2161,6 +2166,12 @@ static int blk_revalidate_zone_cond(struct blk_zone *zone, unsigned int idx, { enum blk_zone_cond cond = zone->cond; + if (idx >= args->nr_zones) { + pr_warn("%s: Zone report index %u exceeds zone count %u\n", + args->disk->disk_name, idx, args->nr_zones); + return -EINVAL; + } + /* Check that the zone condition is consistent with the zone type. */ switch (cond) { case BLK_ZONE_COND_NOT_WP: diff --git a/drivers/accel/ivpu/ivpu_drv.c b/drivers/accel/ivpu/ivpu_drv.c index 35e506074d5f..8c1c87e69f91 100644 --- a/drivers/accel/ivpu/ivpu_drv.c +++ b/drivers/accel/ivpu/ivpu_drv.c @@ -307,6 +307,11 @@ static int ivpu_open(struct drm_device *dev, struct drm_file *file) return -ENODEV; limits = ivpu_user_limits_get(vdev); + if (IS_ERR(limits) && PTR_ERR(limits) == -EMFILE) { + /* Context limit may be held by jobs pending deferred cleanup */ + flush_work(&vdev->job_destroy_work); + limits = ivpu_user_limits_get(vdev); + } if (IS_ERR(limits)) { ret = PTR_ERR(limits); goto err_dev_exit; @@ -510,9 +515,10 @@ void ivpu_prepare_for_reset(struct ivpu_device *vdev) { ivpu_hw_irq_disable(vdev); disable_irq(vdev->irq); - flush_work(&vdev->irq_ipc_work); + atomic_set(&vdev->job_timeout_detected, 0); flush_work(&vdev->irq_dct_work); flush_work(&vdev->context_abort_work); + flush_work(&vdev->job_destroy_work); ivpu_ipc_disable(vdev); ivpu_mmu_disable(vdev); } @@ -584,6 +590,11 @@ static const struct drm_driver driver = { .major = 1, }; +static void ivpu_destroy_workqueue(void *wq) +{ + destroy_workqueue(wq); +} + static int ivpu_irq_init(struct ivpu_device *vdev) { struct pci_dev *pdev = to_pci_dev(vdev->drm.dev); @@ -595,16 +606,26 @@ static int ivpu_irq_init(struct ivpu_device *vdev) return ret; } - INIT_WORK(&vdev->irq_ipc_work, ivpu_ipc_irq_work_fn); INIT_WORK(&vdev->irq_dct_work, ivpu_pm_irq_dct_work_fn); INIT_WORK(&vdev->context_abort_work, ivpu_context_abort_work_fn); + init_llist_head(&vdev->job_destroy_list); + INIT_WORK(&vdev->job_destroy_work, ivpu_job_destroy_work_fn); + + vdev->job_destroy_wq = alloc_workqueue("ivpu_job_destroy", WQ_UNBOUND | WQ_MEM_RECLAIM, 0); + if (!vdev->job_destroy_wq) + return -ENOMEM; + + ret = devm_add_action_or_reset(vdev->drm.dev, ivpu_destroy_workqueue, vdev->job_destroy_wq); + if (ret) + return ret; ivpu_irq_handlers_init(vdev); vdev->irq = pci_irq_vector(pdev, 0); - ret = devm_request_irq(vdev->drm.dev, vdev->irq, ivpu_hw_irq_handler, - IRQF_NO_AUTOEN, DRIVER_NAME, vdev); + ret = devm_request_threaded_irq(vdev->drm.dev, vdev->irq, ivpu_hw_irq_handler, + ivpu_ipc_irq_thread_handler, IRQF_NO_AUTOEN, + DRIVER_NAME, vdev); if (ret) ivpu_err(vdev, "Failed to request an IRQ %d\n", ret); @@ -690,7 +711,7 @@ static int ivpu_dev_init(struct ivpu_device *vdev) vdev->context_xa_limit.max = IVPU_USER_CONTEXT_MAX_SSID; atomic64_set(&vdev->unique_id_counter, 0); atomic_set(&vdev->job_timeout_counter, 0); - atomic_set(&vdev->faults_detected, 0); + atomic_set(&vdev->job_timeout_detected, 0); xa_init_flags(&vdev->context_xa, XA_FLAGS_ALLOC | XA_FLAGS_LOCK_IRQ); xa_init_flags(&vdev->submitted_jobs_xa, XA_FLAGS_ALLOC1); xa_init_flags(&vdev->db_xa, XA_FLAGS_ALLOC1); diff --git a/drivers/accel/ivpu/ivpu_drv.h b/drivers/accel/ivpu/ivpu_drv.h index 9eefbbb7ba11..6f4012926478 100644 --- a/drivers/accel/ivpu/ivpu_drv.h +++ b/drivers/accel/ivpu/ivpu_drv.h @@ -13,6 +13,7 @@ #include #include +#include #include #include #include @@ -157,9 +158,11 @@ struct ivpu_device { struct xa_limit db_limit; u32 db_next; - struct work_struct irq_ipc_work; struct work_struct irq_dct_work; struct work_struct context_abort_work; + struct llist_head job_destroy_list; + struct work_struct job_destroy_work; + struct workqueue_struct *job_destroy_wq; struct mutex bo_list_lock; /* Protects bo_list */ struct list_head bo_list; @@ -168,7 +171,7 @@ struct ivpu_device { struct xarray submitted_jobs_xa; struct ivpu_ipc_consumer job_done_consumer; atomic_t job_timeout_counter; - atomic_t faults_detected; + atomic_t job_timeout_detected; atomic64_t unique_id_counter; diff --git a/drivers/accel/ivpu/ivpu_hw.c b/drivers/accel/ivpu/ivpu_hw.c index d4a9bcda4100..647dc045c231 100644 --- a/drivers/accel/ivpu/ivpu_hw.c +++ b/drivers/accel/ivpu/ivpu_hw.c @@ -399,6 +399,10 @@ irqreturn_t ivpu_hw_irq_handler(int irq, void *ptr) return IRQ_NONE; pm_runtime_mark_last_busy(vdev->drm.dev); + + if (ip_handled) + return IRQ_WAKE_THREAD; + return IRQ_HANDLED; } diff --git a/drivers/accel/ivpu/ivpu_ipc.c b/drivers/accel/ivpu/ivpu_ipc.c index 8f69fb133e2e..459e740873ab 100644 --- a/drivers/accel/ivpu/ivpu_ipc.c +++ b/drivers/accel/ivpu/ivpu_ipc.c @@ -463,13 +463,11 @@ void ivpu_ipc_irq_handler(struct ivpu_device *vdev) ivpu_ipc_rx_mark_free(vdev, ipc_hdr, jsm_msg); } } - - queue_work(system_percpu_wq, &vdev->irq_ipc_work); } -void ivpu_ipc_irq_work_fn(struct work_struct *work) +irqreturn_t ivpu_ipc_irq_thread_handler(int irq, void *ptr) { - struct ivpu_device *vdev = container_of(work, struct ivpu_device, irq_ipc_work); + struct ivpu_device *vdev = ptr; struct ivpu_ipc_info *ipc = vdev->ipc; struct ivpu_ipc_rx_msg *rx_msg, *r; struct list_head cb_msg_list; @@ -484,6 +482,8 @@ void ivpu_ipc_irq_work_fn(struct work_struct *work) rx_msg->callback(vdev, rx_msg->ipc_hdr, rx_msg->jsm_msg); ivpu_ipc_rx_msg_del(vdev, rx_msg); } + + return IRQ_HANDLED; } int ivpu_ipc_init(struct ivpu_device *vdev) diff --git a/drivers/accel/ivpu/ivpu_ipc.h b/drivers/accel/ivpu/ivpu_ipc.h index b524a1985b9d..1c9cf67f1e5b 100644 --- a/drivers/accel/ivpu/ivpu_ipc.h +++ b/drivers/accel/ivpu/ivpu_ipc.h @@ -90,7 +90,7 @@ void ivpu_ipc_disable(struct ivpu_device *vdev); void ivpu_ipc_reset(struct ivpu_device *vdev); void ivpu_ipc_irq_handler(struct ivpu_device *vdev); -void ivpu_ipc_irq_work_fn(struct work_struct *work); +irqreturn_t ivpu_ipc_irq_thread_handler(int irq, void *ptr); void ivpu_ipc_consumer_add(struct ivpu_device *vdev, struct ivpu_ipc_consumer *cons, u32 channel, ivpu_ipc_rx_callback_t callback); diff --git a/drivers/accel/ivpu/ivpu_job.c b/drivers/accel/ivpu/ivpu_job.c index b24f31a8b567..4689b8ab519d 100644 --- a/drivers/accel/ivpu/ivpu_job.c +++ b/drivers/accel/ivpu/ivpu_job.c @@ -535,6 +535,20 @@ static void ivpu_job_destroy(struct ivpu_job *job) kfree(job); } +void ivpu_job_destroy_work_fn(struct work_struct *work) +{ + struct ivpu_device *vdev = container_of(work, struct ivpu_device, job_destroy_work); + struct ivpu_job *job, *tmp; + struct llist_node *list; + + list = llist_del_all(&vdev->job_destroy_list); + + llist_for_each_entry_safe(job, tmp, list, destroy_node) { + ivpu_job_destroy(job); + ivpu_rpm_put(vdev); + } +} + static struct ivpu_job * ivpu_job_create(struct ivpu_file_priv *file_priv, u32 engine_idx, u32 bo_count) { @@ -607,7 +621,6 @@ bool ivpu_job_handle_engine_error(struct ivpu_device *vdev, u32 job_id, u32 job_ * status and ensure both are handled in the same way */ job->file_priv->has_mmu_faults = true; - atomic_set(&vdev->faults_detected, 1); queue_work(system_percpu_wq, &vdev->context_abort_work); return true; } @@ -619,7 +632,7 @@ bool ivpu_job_handle_engine_error(struct ivpu_device *vdev, u32 job_id, u32 job_ return false; } -static int ivpu_job_signal_and_destroy(struct ivpu_device *vdev, u32 job_id, u32 job_status) +static struct ivpu_job *ivpu_job_signal(struct ivpu_device *vdev, u32 job_id, u32 job_status) { struct ivpu_job *job; @@ -627,7 +640,7 @@ static int ivpu_job_signal_and_destroy(struct ivpu_device *vdev, u32 job_id, u32 job = xa_load(&vdev->submitted_jobs_xa, job_id); if (!job) - return -ENOENT; + return NULL; ivpu_job_remove_from_submitted_jobs(vdev, job_id); @@ -646,14 +659,37 @@ static int ivpu_job_signal_and_destroy(struct ivpu_device *vdev, u32 job_id, u32 job->job_id, job->file_priv->ctx.id, job->cmdq_id, job->engine_idx, job->job_status); - ivpu_job_destroy(job); ivpu_stop_job_timeout_detection(vdev); - ivpu_rpm_put(vdev); - if (!xa_empty(&vdev->submitted_jobs_xa)) ivpu_start_job_timeout_detection(vdev); + return job; +} + +static int ivpu_job_signal_and_destroy(struct ivpu_device *vdev, u32 job_id, u32 job_status) +{ + struct ivpu_job *job = ivpu_job_signal(vdev, job_id, job_status); + + if (!job) + return -ENOENT; + + ivpu_job_destroy(job); + ivpu_rpm_put(vdev); + + return 0; +} + +static int ivpu_job_signal_and_defer_destroy(struct ivpu_device *vdev, u32 job_id, u32 job_status) +{ + struct ivpu_job *job = ivpu_job_signal(vdev, job_id, job_status); + + if (!job) + return -ENOENT; + + llist_add(&job->destroy_node, &vdev->job_destroy_list); + queue_work(vdev->job_destroy_wq, &vdev->job_destroy_work); + return 0; } @@ -689,6 +725,7 @@ static int ivpu_job_submit(struct ivpu_job *job, u8 priority, u32 cmdq_id) struct ivpu_file_priv *file_priv = job->file_priv; struct ivpu_device *vdev = job->vdev; struct ivpu_cmdq *cmdq; + bool flushed = false; bool is_first_job; int ret; @@ -696,6 +733,7 @@ static int ivpu_job_submit(struct ivpu_job *job, u8 priority, u32 cmdq_id) if (ret < 0) return ret; +retry: mutex_lock(&vdev->submitted_jobs_lock); mutex_lock(&file_priv->lock); @@ -709,6 +747,14 @@ static int ivpu_job_submit(struct ivpu_job *job, u8 priority, u32 cmdq_id) } ret = ivpu_cmdq_register(file_priv, cmdq); + if (ret == -EBUSY && !flushed) { + /* Doorbell may be held by jobs pending deferred cleanup */ + mutex_unlock(&file_priv->lock); + mutex_unlock(&vdev->submitted_jobs_lock); + flush_work(&vdev->job_destroy_work); + flushed = true; + goto retry; + } if (ret) { ivpu_err(vdev, "Failed to register command queue: %d\n", ret); goto err_unlock; @@ -1101,7 +1147,7 @@ ivpu_job_done_callback(struct ivpu_device *vdev, struct ivpu_ipc_hdr *ipc_hdr, mutex_lock(&vdev->submitted_jobs_lock); if (!ivpu_job_handle_engine_error(vdev, payload->job_id, payload->job_status)) /* No engine error, complete the job normally */ - ivpu_job_signal_and_destroy(vdev, payload->job_id, payload->job_status); + ivpu_job_signal_and_defer_destroy(vdev, payload->job_id, payload->job_status); mutex_unlock(&vdev->submitted_jobs_lock); } @@ -1128,10 +1174,10 @@ static int reset_engine_and_mark_faulty_contexts(struct ivpu_device *vdev) return ret; /* - * If faults are detected, ignore guilty contexts from engine reset as NPU may not be stuck - * and could return currently running good context and faulty contexts are already marked + * If job timeout is detected, read guilty context from engine reset, for other reasons + * faulty context is already known */ - if (atomic_cmpxchg(&vdev->faults_detected, 1, 0) == 1) + if (atomic_cmpxchg(&vdev->job_timeout_detected, 1, 0) == 0) return 0; num_impacted_contexts = resp.payload.engine_reset_done.num_impacted_contexts; diff --git a/drivers/accel/ivpu/ivpu_job.h b/drivers/accel/ivpu/ivpu_job.h index 3ab61e6a5616..d8dbce82447a 100644 --- a/drivers/accel/ivpu/ivpu_job.h +++ b/drivers/accel/ivpu/ivpu_job.h @@ -6,8 +6,10 @@ #ifndef __IVPU_JOB_H__ #define __IVPU_JOB_H__ -#include #include +#include +#include +#include #include "ivpu_gem.h" @@ -47,6 +49,7 @@ struct ivpu_cmdq { * @vdev: Pointer to the VPU device * @file_priv: The client context that submitted this job * @done_fence: Fence signaled when job completes + * @destroy_node: List node for deferred resource cleanup after job completion * @cmd_buf_vpu_addr: VPU address of the command buffer for this job * @cmdq_id: Command queue ID used for submission * @job_id: Unique job ID for tracking and status reporting @@ -61,6 +64,7 @@ struct ivpu_job { struct ivpu_device *vdev; struct ivpu_file_priv *file_priv; struct dma_fence *done_fence; + struct llist_node destroy_node; u64 cmd_buf_vpu_addr; u32 cmdq_id; u32 job_id; @@ -87,6 +91,7 @@ void ivpu_job_done_consumer_init(struct ivpu_device *vdev); void ivpu_job_done_consumer_fini(struct ivpu_device *vdev); bool ivpu_job_handle_engine_error(struct ivpu_device *vdev, u32 job_id, u32 job_status); void ivpu_context_abort_work_fn(struct work_struct *work); +void ivpu_job_destroy_work_fn(struct work_struct *work); void ivpu_jobs_abort_all(struct ivpu_device *vdev); diff --git a/drivers/accel/ivpu/ivpu_mmu.c b/drivers/accel/ivpu/ivpu_mmu.c index 41efd8985fa6..b2025274f91d 100644 --- a/drivers/accel/ivpu/ivpu_mmu.c +++ b/drivers/accel/ivpu/ivpu_mmu.c @@ -964,7 +964,6 @@ void ivpu_mmu_irq_evtq_handler(struct ivpu_device *vdev) file_priv = xa_load(&vdev->context_xa, ssid); if (file_priv) { if (!READ_ONCE(file_priv->has_mmu_faults)) { - atomic_set(&vdev->faults_detected, 1); ivpu_mmu_dump_event(vdev, event); WRITE_ONCE(file_priv->has_mmu_faults, true); } diff --git a/drivers/accel/ivpu/ivpu_pm.c b/drivers/accel/ivpu/ivpu_pm.c index c1ce8329790e..de0becbfdffb 100644 --- a/drivers/accel/ivpu/ivpu_pm.c +++ b/drivers/accel/ivpu/ivpu_pm.c @@ -229,6 +229,7 @@ static void ivpu_job_timeout_work(struct work_struct *work) ivpu_jsm_state_dump(vdev); ivpu_dev_coredump(vdev); + atomic_set(&vdev->job_timeout_detected, 1); queue_work(system_percpu_wq, &vdev->context_abort_work); } diff --git a/drivers/ata/libata-scsi.c b/drivers/ata/libata-scsi.c index f0bb4068d14d..3559485b5034 100644 --- a/drivers/ata/libata-scsi.c +++ b/drivers/ata/libata-scsi.c @@ -261,12 +261,18 @@ static void ata_scsi_set_passthru_sense_fields(struct ata_queued_cmd *qc) /* descriptor format */ len = sb[7]; - desc = (char *)scsi_sense_desc_find(sb, len + 8, 9); + desc = (char *)scsi_sense_desc_find(sb, SCSI_SENSE_BUFFERSIZE, 9); if (!desc) { - if (SCSI_SENSE_BUFFERSIZE < len + 14) + /* + * The descriptor is written at sb[8 + len] and is 14 + * bytes long, so it needs len + 22 bytes of buffer. + */ + if (len + 22 > SCSI_SENSE_BUFFERSIZE) return; sb[7] = len + 14; desc = sb + 8 + len; + } else if (desc - sb > SCSI_SENSE_BUFFERSIZE - 14) { + return; } desc[0] = 9; desc[1] = 12; diff --git a/drivers/base/cacheinfo.c b/drivers/base/cacheinfo.c index 9f9c72727a05..7a47a392568a 100644 --- a/drivers/base/cacheinfo.c +++ b/drivers/base/cacheinfo.c @@ -1040,9 +1040,10 @@ static int cacheinfo_cpu_online(unsigned int cpu) rc = cache_add_dev(cpu); if (rc) goto err; - if (cpu_map_shared_cache(true, cpu, &cpu_map)) + if (cpu_map_shared_cache(true, cpu, &cpu_map)) { update_per_cpu_data_slice_size(true, cpu, cpu_map); - sched_update_llc_bytes(cpu); + sched_update_llc_bytes(cpu_map); + } return 0; err: free_cache_attributes(cpu); @@ -1059,10 +1060,10 @@ static int cacheinfo_cpu_pre_down(unsigned int cpu) cpu_cache_sysfs_exit(cpu); free_cache_attributes(cpu); - if (nr_shared > 1) + if (nr_shared > 1) { update_per_cpu_data_slice_size(false, cpu, cpu_map); - - sched_update_llc_bytes(cpu); + sched_update_llc_bytes(cpu_map); + } return 0; } diff --git a/drivers/bluetooth/btintel_pcie.c b/drivers/bluetooth/btintel_pcie.c index 3262e950a6bd..9381895dc3cf 100644 --- a/drivers/bluetooth/btintel_pcie.c +++ b/drivers/bluetooth/btintel_pcie.c @@ -1099,6 +1099,11 @@ static void btintel_pcie_msix_tx_handle(struct btintel_pcie_data *data) txq = &data->txq; + if (cr_hia >= txq->count) { + bt_dev_err(data->hdev, "TXQ: invalid cr_hia %u", cr_hia); + return; + } + while (cr_tia != cr_hia) { data->tx_wait_done = true; wake_up(&data->tx_wait_q); @@ -1588,6 +1593,11 @@ static void btintel_pcie_msix_rx_handle(struct btintel_pcie_data *data) rxq = &data->rxq; + if (cr_hia >= rxq->count) { + bt_dev_err(hdev, "RXQ: invalid cr_hia %u", cr_hia); + return; + } + /* The firmware sends multiple CD in a single MSI-X and it needs to * process all received CDs in this interrupt. */ @@ -1595,6 +1605,12 @@ static void btintel_pcie_msix_rx_handle(struct btintel_pcie_data *data) urbd1 = &rxq->urbd1s[cr_tia]; ipc_print_urbd1(data->hdev, urbd1, cr_tia); + if (urbd1->frbd_tag >= rxq->count) { + bt_dev_err(hdev, "RXQ: invalid frbd_tag %u", + urbd1->frbd_tag); + return; + } + buf = &rxq->bufs[urbd1->frbd_tag]; if (!buf) { bt_dev_err(hdev, "RXQ: failed to get the DMA buffer for %d", diff --git a/drivers/bluetooth/btnxpuart.c b/drivers/bluetooth/btnxpuart.c index e6c15bc6a30b..73b319c7e176 100644 --- a/drivers/bluetooth/btnxpuart.c +++ b/drivers/bluetooth/btnxpuart.c @@ -1397,9 +1397,11 @@ static int nxp_process_fw_dump(struct hci_dev *hdev, struct sk_buff *skb) msecs_to_jiffies(20000)); } - err = hci_devcd_append(hdev, skb_clone(skb, GFP_ATOMIC)); - if (err < 0) - goto free_skb; + if (IS_ENABLED(CONFIG_DEV_COREDUMP)) { + err = hci_devcd_append(hdev, skb_clone(skb, GFP_ATOMIC)); + if (err < 0) + goto free_skb; + } if (buf_len == 0) { bt_dev_warn(hdev, "==== FW dump complete ==="); diff --git a/drivers/dpll/dpll_netlink.c b/drivers/dpll/dpll_netlink.c index 9e55745e33e4..9f274c6253c6 100644 --- a/drivers/dpll/dpll_netlink.c +++ b/drivers/dpll/dpll_netlink.c @@ -1282,8 +1282,7 @@ dpll_pin_ref_sync_state_set(struct dpll_pin *pin, unsigned long i; int ret; - ref_sync_pin = xa_find(&pin->ref_sync_pins, &ref_sync_pin_idx, - ULONG_MAX, XA_PRESENT); + ref_sync_pin = xa_load(&pin->ref_sync_pins, ref_sync_pin_idx); if (!ref_sync_pin) { NL_SET_ERR_MSG(extack, "reference sync pin not found"); return -EINVAL; diff --git a/drivers/firewire/core-cdev.c b/drivers/firewire/core-cdev.c index e49d8a58be09..664952a67a11 100644 --- a/drivers/firewire/core-cdev.c +++ b/drivers/firewire/core-cdev.c @@ -1397,8 +1397,10 @@ static void iso_resource_auto_work(struct work_struct *work) } else { // Transit from allocation to reallocation, except if the client requested // deallocation in the meantime. - scoped_guard(spinlock_irq, &client->lock) - r->todo = ISO_RES_AUTO_REALLOC; + scoped_guard(spinlock_irq, &client->lock) { + if (r->todo == ISO_RES_AUTO_ALLOC) + r->todo = ISO_RES_AUTO_REALLOC; + } if (channel >= 0) r->params.channels_mask = BIT_ULL(channel); diff --git a/drivers/gpio/gpio-arizona.c b/drivers/gpio/gpio-arizona.c index a7e98d395d8e..88cf013f0d90 100644 --- a/drivers/gpio/gpio-arizona.c +++ b/drivers/gpio/gpio-arizona.c @@ -115,8 +115,12 @@ static int arizona_gpio_direction_out(struct gpio_chip *chip, if (value) value = ARIZONA_GPN_LVL; - return regmap_update_bits(arizona->regmap, ARIZONA_GPIO1_CTRL + offset, - ARIZONA_GPN_DIR | ARIZONA_GPN_LVL, value); + ret = regmap_update_bits(arizona->regmap, ARIZONA_GPIO1_CTRL + offset, + ARIZONA_GPN_DIR | ARIZONA_GPN_LVL, value); + if (ret < 0 && (val & ARIZONA_GPN_DIR) && persistent) + pm_runtime_put_autosuspend(chip->parent); + + return ret; } static int arizona_gpio_set(struct gpio_chip *chip, unsigned int offset, diff --git a/drivers/gpio/gpio-tps65219.c b/drivers/gpio/gpio-tps65219.c index 457fd8a589e8..6958466455d0 100644 --- a/drivers/gpio/gpio-tps65219.c +++ b/drivers/gpio/gpio-tps65219.c @@ -79,7 +79,7 @@ static int tps65219_gpio_get(struct gpio_chip *gc, unsigned int offset) if (ret) return ret; - ret = !!(val & BIT(TPS65219_MFP_GPIO_STATUS_MASK)); + ret = !!(val & TPS65219_MFP_GPIO_STATUS_MASK); dev_warn(dev, "GPIO%d = %d, MULTI_DEVICE_ENABLE, not a standard GPIO\n", offset, ret); /* @@ -87,7 +87,7 @@ static int tps65219_gpio_get(struct gpio_chip *gc, unsigned int offset) * status bit. */ - if (tps65219_gpio_get_direction(gc, offset) == GPIO_LINE_DIRECTION_OUT) + if (gc->get_direction(gc, offset) == GPIO_LINE_DIRECTION_OUT) return -ENOTSUPP; return ret; @@ -158,8 +158,10 @@ static int tps65214_gpio_change_direction(struct gpio_chip *gc, unsigned int off if (ret) dev_err(dev, "GPIO%d configured as VSEL, not GPIO\n", offset); + val = direction == GPIO_LINE_DIRECTION_OUT ? + TPS65214_GPIO0_DIR_MASK : 0; ret = regmap_update_bits(gpio->tps->regmap, TPS65219_REG_GENERAL_CONFIG, - TPS65214_GPIO0_DIR_MASK, direction); + TPS65214_GPIO0_DIR_MASK, val); if (ret) dev_err(dev, "Fail to change direction to %u for GPIO%d.\n", direction, offset); @@ -176,7 +178,7 @@ static int tps65219_gpio_direction_input(struct gpio_chip *gc, unsigned int offs return -ENOTSUPP; } - if (tps65219_gpio_get_direction(gc, offset) == GPIO_LINE_DIRECTION_IN) + if (gc->get_direction(gc, offset) == GPIO_LINE_DIRECTION_IN) return 0; return gpio->change_dir(gc, offset, GPIO_LINE_DIRECTION_IN); @@ -190,7 +192,7 @@ static int tps65219_gpio_direction_output(struct gpio_chip *gc, unsigned int off if (offset != TPS6521X_GPIO0_IDX) return 0; - if (tps65219_gpio_get_direction(gc, offset) == GPIO_LINE_DIRECTION_OUT) + if (gc->get_direction(gc, offset) == GPIO_LINE_DIRECTION_OUT) return 0; return gpio->change_dir(gc, offset, GPIO_LINE_DIRECTION_OUT); diff --git a/drivers/gpio/gpio-zynq.c b/drivers/gpio/gpio-zynq.c index 15a79d9a2e9e..13e4c5f0e297 100644 --- a/drivers/gpio/gpio-zynq.c +++ b/drivers/gpio/gpio-zynq.c @@ -798,15 +798,7 @@ static int zynq_gpio_runtime_resume(struct device *dev) static int zynq_gpio_request(struct gpio_chip *chip, unsigned int offset) { - int ret; - - ret = pm_runtime_get_sync(chip->parent); - - /* - * If the device is already active pm_runtime_get() will return 1 on - * success, but gpio_request still needs to return 0. - */ - return ret < 0 ? ret : 0; + return pm_runtime_resume_and_get(chip->parent); } static void zynq_gpio_free(struct gpio_chip *chip, unsigned int offset) diff --git a/drivers/gpio/gpiolib-cdev.c b/drivers/gpio/gpiolib-cdev.c index 82f27db0b230..4d0a94e19eba 100644 --- a/drivers/gpio/gpiolib-cdev.c +++ b/drivers/gpio/gpiolib-cdev.c @@ -2172,18 +2172,19 @@ static void gpio_v2_line_info_changed_to_v1( #endif /* CONFIG_GPIO_CDEV_V1 */ -static void gpio_desc_to_lineinfo(struct gpio_desc *desc, - struct gpio_v2_line_info *info, bool atomic) +static int gpio_desc_to_lineinfo(struct gpio_desc *desc, + struct gpio_v2_line_info *info, bool atomic) { u32 debounce_period_us; unsigned long dflags; const char *label; + memset(info, 0, sizeof(*info)); + CLASS(gpio_chip_guard, guard)(desc); if (!guard.gc) - return; + return -ENODEV; - memset(info, 0, sizeof(*info)); info->offset = gpiod_hwgpio(desc); if (desc->name) @@ -2258,6 +2259,8 @@ static void gpio_desc_to_lineinfo(struct gpio_desc *desc, debounce_period_us; info->num_attrs++; } + + return 0; } struct gpio_chardev_data { @@ -2309,6 +2312,7 @@ static int lineinfo_get_v1(struct gpio_chardev_data *cdev, void __user *ip, struct gpio_desc *desc; struct gpioline_info lineinfo; struct gpio_v2_line_info lineinfo_v2; + int ret; if (copy_from_user(&lineinfo, ip, sizeof(lineinfo))) return -EFAULT; @@ -2326,7 +2330,10 @@ static int lineinfo_get_v1(struct gpio_chardev_data *cdev, void __user *ip, return -EBUSY; } - gpio_desc_to_lineinfo(desc, &lineinfo_v2, false); + ret = gpio_desc_to_lineinfo(desc, &lineinfo_v2, false); + if (ret) + return ret; + gpio_v2_line_info_to_v1(&lineinfo_v2, &lineinfo); if (copy_to_user(ip, &lineinfo, sizeof(lineinfo))) { @@ -2344,6 +2351,7 @@ static int lineinfo_get(struct gpio_chardev_data *cdev, void __user *ip, { struct gpio_desc *desc; struct gpio_v2_line_info lineinfo; + int ret; if (copy_from_user(&lineinfo, ip, sizeof(lineinfo))) return -EFAULT; @@ -2363,7 +2371,10 @@ static int lineinfo_get(struct gpio_chardev_data *cdev, void __user *ip, if (test_and_set_bit(lineinfo.offset, cdev->watched_lines)) return -EBUSY; } - gpio_desc_to_lineinfo(desc, &lineinfo, false); + + ret = gpio_desc_to_lineinfo(desc, &lineinfo, false); + if (ret) + return ret; if (copy_to_user(ip, &lineinfo, sizeof(lineinfo))) { if (watch) @@ -2489,6 +2500,7 @@ static int lineinfo_changed_notify(struct notifier_block *nb, struct lineinfo_changed_ctx *ctx; struct gpio_desc *desc = data; struct file *fp; + int ret; if (!test_bit(gpiod_hwgpio(desc), cdev->watched_lines)) return NOTIFY_DONE; @@ -2519,7 +2531,13 @@ static int lineinfo_changed_notify(struct notifier_block *nb, ctx->chg.event_type = action; ctx->chg.timestamp_ns = ktime_get_ns(); - gpio_desc_to_lineinfo(desc, &ctx->chg.info, true); + + ret = gpio_desc_to_lineinfo(desc, &ctx->chg.info, true); + if (ret) { + fput(fp); + return NOTIFY_DONE; + } + /* Keep the GPIO device alive until we emit the event. */ ctx->gdev = gpio_device_get(desc->gdev); ctx->cdev = cdev; diff --git a/drivers/gpio/gpiolib.c b/drivers/gpio/gpiolib.c index ef8ccaf17c9c..28f7265c5d56 100644 --- a/drivers/gpio/gpiolib.c +++ b/drivers/gpio/gpiolib.c @@ -1029,6 +1029,13 @@ int gpiochip_add_hog(struct gpio_chip *gc, struct fwnode_handle *fwnode) ret = of_gpiochip_get_lflags(gc, &gpiospec, &lflags); if (ret) return ret; + + /* + * If no line-name property is present, fall back to the OF + * node name as in the previous implementation. + */ + if (!name) + name = to_of_node(fwnode)->name; } else { /* * GPIO_ACTIVE_LOW is currently the only lookup flag diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_acpi.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_acpi.c index 516ab9cf88fc..8d3f1f5a141e 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_acpi.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_acpi.c @@ -1167,8 +1167,10 @@ int amdgpu_acpi_enumerate_xcc(void) } xcc_info = kzalloc_obj(struct amdgpu_acpi_xcc_info); - if (!xcc_info) + if (!xcc_info) { + acpi_dev_put(acpi_dev); return -ENOMEM; + } INIT_LIST_HEAD(&xcc_info->list); xcc_info->handle = acpi_device_handle(acpi_dev); diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_debugfs.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_debugfs.c index df2588642817..4536d46b7f21 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_debugfs.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_debugfs.c @@ -1773,8 +1773,10 @@ static int amdgpu_debugfs_test_ib_show(struct seq_file *m, void *unused) /* Avoid accidently unparking the sched thread during GPU reset */ r = down_write_killable(&adev->reset_domain->sem); - if (r) + if (r) { + pm_runtime_put_autosuspend(dev->dev); return r; + } /* hold on the scheduler */ for (i = 0; i < AMDGPU_MAX_RINGS; i++) { diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c index f6b7522c3c82..f8652fd0525d 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_eviction_fence.c @@ -68,6 +68,9 @@ amdgpu_eviction_fence_suspend_worker(struct work_struct *work) mutex_lock(&uq_mgr->userq_mutex); + /* Fence waits are not allowed in a fence signalling critical section. */ + amdgpu_userq_wait_for_signal(uq_mgr); + /* * This is intentionally after taking the userq_mutex since we do * allocate memory while holding this lock, but only after ensuring that diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.c index b97fa35bac23..4d5d95e74bb5 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_ring.c @@ -254,7 +254,6 @@ int amdgpu_ring_init(struct amdgpu_device *adev, struct amdgpu_ring *ring, ring->adev = adev; ring->num_hw_submission = sched_hw_submission; ring->sched_score = sched_score; - ring->vmid_wait = dma_fence_get_stub(); ring->idx = adev->num_rings++; adev->rings[ring->idx] = ring; @@ -374,6 +373,7 @@ int amdgpu_ring_init(struct amdgpu_device *adev, struct amdgpu_ring *ring, ring->max_dw = max_dw; ring->hw_prio = hw_prio; + ring->vmid_wait = dma_fence_get_stub(); if (!ring->no_scheduler && ring->funcs->type < AMDGPU_HW_IP_NUM) { hw_ip = ring->funcs->type; diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c index b18d78720656..655429c40351 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.c @@ -167,27 +167,27 @@ static void amdgpu_userq_hang_detect_work(struct work_struct *work) void amdgpu_userq_start_hang_detect_work(struct amdgpu_usermode_queue *queue) { struct amdgpu_device *adev; - unsigned long timeout_ms; + unsigned long timeout_jiffies; adev = queue->userq_mgr->adev; /* Determine timeout based on queue type */ switch (queue->queue_type) { case AMDGPU_RING_TYPE_GFX: - timeout_ms = adev->gfx_timeout; + timeout_jiffies = adev->gfx_timeout; break; case AMDGPU_RING_TYPE_COMPUTE: - timeout_ms = adev->compute_timeout; + timeout_jiffies = adev->compute_timeout; break; case AMDGPU_RING_TYPE_SDMA: - timeout_ms = adev->sdma_timeout; + timeout_jiffies = adev->sdma_timeout; break; default: - timeout_ms = adev->gfx_timeout; + timeout_jiffies = adev->gfx_timeout; break; } queue_delayed_work(adev->reset_domain->wq, &queue->hang_detect_work, - msecs_to_jiffies(timeout_ms)); + timeout_jiffies); } void amdgpu_userq_process_fence_irq(struct amdgpu_device *adev, u32 doorbell) @@ -1168,7 +1168,7 @@ amdgpu_userq_evict_all(struct amdgpu_userq_mgr *uq_mgr) return ret; } -static void +void amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr) { struct amdgpu_usermode_queue *queue; @@ -1187,8 +1187,6 @@ amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr) void amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr) { - /* Wait for any pending userqueue fence work to finish */ - amdgpu_userq_wait_for_signal(uq_mgr); amdgpu_userq_evict_all(uq_mgr); } diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h index d1751febaefe..ff9f203a320b 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_userq.h @@ -156,6 +156,7 @@ void amdgpu_userq_mgr_cancel_reset_work(struct amdgpu_device *adev); void amdgpu_userq_mgr_cancel_resume(struct amdgpu_userq_mgr *userq_mgr); void amdgpu_userq_mgr_fini(struct amdgpu_userq_mgr *userq_mgr); +void amdgpu_userq_wait_for_signal(struct amdgpu_userq_mgr *uq_mgr); void amdgpu_userq_evict(struct amdgpu_userq_mgr *uq_mgr); void amdgpu_userq_ensure_ev_fence(struct amdgpu_userq_mgr *userq_mgr, diff --git a/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c b/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c index 816800f596cd..5edf00a8273b 100644 --- a/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c +++ b/drivers/gpu/drm/amd/amdgpu/amdgpu_vm.c @@ -2665,6 +2665,7 @@ int amdgpu_vm_init(struct amdgpu_device *adev, struct amdgpu_vm *vm, amdgpu_bo_unref(&root_bo); error_free_delayed: + dma_fence_put(vm->last_update); dma_fence_put(vm->last_tlb_flush); dma_fence_put(vm->last_unlocked); ttm_lru_bulk_move_fini(&adev->mman.bdev, &vm->lru_bulk_move); diff --git a/drivers/gpu/drm/amd/amdgpu/vcn_v4_0_3.c b/drivers/gpu/drm/amd/amdgpu/vcn_v4_0_3.c index 7f001c32e911..55acb544b1b6 100644 --- a/drivers/gpu/drm/amd/amdgpu/vcn_v4_0_3.c +++ b/drivers/gpu/drm/amd/amdgpu/vcn_v4_0_3.c @@ -1672,7 +1672,8 @@ static int vcn_v4_0_3_reset_jpeg_pre_helper(struct amdgpu_device *adev, int inst /* if Jobs are still pending after timeout, * We'll handle them in the bottom helper */ - amdgpu_fence_wait_polling(ring, wait_seq, adev->video_timeout); + amdgpu_fence_wait_polling(ring, wait_seq, + jiffies_to_usecs(adev->video_timeout)); } return 0; diff --git a/drivers/gpu/drm/amd/amdgpu/vcn_v5_0_1.c b/drivers/gpu/drm/amd/amdgpu/vcn_v5_0_1.c index d3db0494341e..d25a775b3efd 100644 --- a/drivers/gpu/drm/amd/amdgpu/vcn_v5_0_1.c +++ b/drivers/gpu/drm/amd/amdgpu/vcn_v5_0_1.c @@ -1318,7 +1318,8 @@ static int vcn_v5_0_1_reset_jpeg_pre_helper(struct amdgpu_device *adev, int inst /* if Jobs are still pending after timeout, * We'll handle them in the bottom helper */ - amdgpu_fence_wait_polling(ring, wait_seq, adev->video_timeout); + amdgpu_fence_wait_polling(ring, wait_seq, + jiffies_to_usecs(adev->video_timeout)); } return 0; diff --git a/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c b/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c index 6738cd1d4f15..2e85a1de51c0 100644 --- a/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c +++ b/drivers/gpu/drm/amd/amdkfd/kfd_chardev.c @@ -35,6 +35,7 @@ #include #include #include +#include #include #include #include @@ -70,18 +71,54 @@ static const struct class kfd_class = { }; /* - * Cache the address space of the chardev on first open so that the reset - * path can drop all userspace mappings of doorbell and MMIO ranges via - * unmap_mapping_range(). + * Private pseudo-filesystem for KFD, Provides a stable, module-owned + * inode whose address_space is the unmap target for all /dev/kfd + * openers during GPU reset. */ -static struct address_space *kfd_dev_mapping; +static struct vfsmount *kfd_fs_mnt; +static int kfd_fs_cnt; -void kfd_dev_unmap_mapping_range(loff_t const holebegin, loff_t const holelen) +static int kfd_fs_init_fs_context(struct fs_context *fc) +{ + return init_pseudo(fc, 0x4b464400 /* "KFD" */) ? 0 : -ENOMEM; +} + +static struct file_system_type kfd_fs_type = { + .name = "kfd", + .init_fs_context = kfd_fs_init_fs_context, + .kill_sb = kill_anon_super, +}; + +static struct inode *kfd_fs_inode_new(void) { - struct address_space *mapping = READ_ONCE(kfd_dev_mapping); + struct inode *inode; + int r; + + r = simple_pin_fs(&kfd_fs_type, &kfd_fs_mnt, &kfd_fs_cnt); + if (r < 0) + return ERR_PTR(r); + + inode = alloc_anon_inode(kfd_fs_mnt->mnt_sb); + if (IS_ERR(inode)) + simple_release_fs(&kfd_fs_mnt, &kfd_fs_cnt); - if (mapping) - unmap_mapping_range(mapping, holebegin, holelen, 1); + return inode; +} + +static void kfd_fs_inode_free(struct inode *inode) +{ + if (inode) { + iput(inode); + simple_release_fs(&kfd_fs_mnt, &kfd_fs_cnt); + } +} + +static struct inode *kfd_anon_inode; + +void kfd_dev_unmap_mapping_range(loff_t const holebegin, loff_t const holelen) +{ + if (kfd_anon_inode) + unmap_mapping_range(kfd_anon_inode->i_mapping, holebegin, holelen, 1); } static inline struct kfd_process_device *kfd_lock_pdd_by_id(struct kfd_process *p, __u32 gpu_id) @@ -107,6 +144,13 @@ int kfd_chardev_init(void) { int err = 0; + kfd_anon_inode = kfd_fs_inode_new(); + if (IS_ERR(kfd_anon_inode)) { + err = PTR_ERR(kfd_anon_inode); + kfd_anon_inode = NULL; + return err; + } + kfd_char_dev_major = register_chrdev(0, kfd_dev_name, &kfd_fops); err = kfd_char_dev_major; if (err < 0) @@ -130,6 +174,8 @@ int kfd_chardev_init(void) err_class_create: unregister_chrdev(kfd_char_dev_major, kfd_dev_name); err_register_chrdev: + kfd_fs_inode_free(kfd_anon_inode); + kfd_anon_inode = NULL; return err; } @@ -138,6 +184,8 @@ void kfd_chardev_exit(void) device_destroy(&kfd_class, MKDEV(kfd_char_dev_major, 0)); class_unregister(&kfd_class); unregister_chrdev(kfd_char_dev_major, kfd_dev_name); + kfd_fs_inode_free(kfd_anon_inode); + kfd_anon_inode = NULL; kfd_device = NULL; } @@ -150,12 +198,7 @@ static int kfd_open(struct inode *inode, struct file *filep) if (iminor(inode) != 0) return -ENODEV; - /* - * /dev/kfd is a single chardev so all opens share one inode. Cache - * its address_space on the first open for use by the reset path. - */ - if (!READ_ONCE(kfd_dev_mapping)) - cmpxchg(&kfd_dev_mapping, NULL, inode->i_mapping); + filep->f_mapping = kfd_anon_inode->i_mapping; is_32bit_user_mode = in_compat_syscall(); diff --git a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c index 5ba196dd9d7a..e5945485ec59 100644 --- a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c +++ b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.c @@ -3344,7 +3344,7 @@ static void dm_gpureset_toggle_interrupts(struct amdgpu_device *adev, if (acrtc && state->stream_status[i].plane_count != 0 && amdgpu_ip_version(adev, DCE_HWIP, 0) == 0) { irq_source = IRQ_TYPE_PFLIP + acrtc->otg_inst; - rc = dc_interrupt_set(adev->dm.dc, irq_source, enable) ? 0 : -EBUSY; + rc = amdgpu_dm_irq_set(adev, irq_source, enable) ? 0 : -EBUSY; if (rc) drm_warn(adev_to_drm(adev), "Failed to %s pflip interrupts\n", enable ? "enable" : "disable"); @@ -3368,7 +3368,7 @@ static void dm_gpureset_toggle_interrupts(struct amdgpu_device *adev, /* During gpu-reset we disable and then enable vblank irq, so * don't use amdgpu_irq_get/put() to avoid refcount change. */ - if (!dc_interrupt_set(adev->dm.dc, irq_source, enable)) + if (!amdgpu_dm_irq_set(adev, irq_source, enable)) drm_warn(adev_to_drm(adev), "Failed to %sable vblank interrupt\n", enable ? "en" : "dis"); } else if (acrtc && state->stream_status[i].plane_count != 0) { @@ -12348,8 +12348,10 @@ static int dm_update_crtc_state(struct amdgpu_display_manager *dm, skip_modeset: /* Release extra reference */ - if (new_stream) + if (new_stream) { dc_stream_release(new_stream); + new_stream = NULL; + } new_stream = NULL; /* diff --git a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.h b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.h index 797f94471810..d2602c0310b2 100644 --- a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.h +++ b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm.h @@ -550,6 +550,18 @@ struct amdgpu_display_manager { struct common_irq_params vupdate_params[DC_IRQ_SOURCE_VUPDATE6 - DC_IRQ_SOURCE_VUPDATE1 + 1]; + /** + * @irq_reg_lock: + * + * Serializes the read-modify-writes of the HW interrupt control + * registers. Several interrupt sources share one register - e.g. the + * enable and clear bits of both VSTARTUP (vblank) and VUPDATE_NO_LOCK + * live in OTG_GLOBAL_SYNC_STATUS. Therefore, enabling one source must + * not race with acking another. Held only across amdgpu_dm_irq_set() + * and amdgpu_dm_irq_ack(). + */ + spinlock_t irq_reg_lock; + /** * @dmub_trace_params: * diff --git a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_crtc.c b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_crtc.c index f47ee9937ada..7206439ea5d5 100644 --- a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_crtc.c +++ b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_crtc.c @@ -31,6 +31,7 @@ #include "amdgpu_dm_psr.h" #include "amdgpu_dm_replay.h" #include "amdgpu_dm_crtc.h" +#include "amdgpu_dm_irq.h" #include "amdgpu_dm_plane.h" #include "amdgpu_dm_trace.h" #include "amdgpu_dm_debugfs.h" @@ -87,7 +88,7 @@ int amdgpu_dm_crtc_set_vupdate_irq(struct drm_crtc *crtc, bool enable) irq_source = IRQ_TYPE_VUPDATE + acrtc->otg_inst; - rc = dc_interrupt_set(adev->dm.dc, irq_source, enable) ? 0 : -EBUSY; + rc = amdgpu_dm_irq_set(adev, irq_source, enable) ? 0 : -EBUSY; DRM_DEBUG_VBL("crtc %d - vupdate irq %sabling: r=%d\n", acrtc->crtc_id, enable ? "en" : "dis", rc); diff --git a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_helpers.c b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_helpers.c index fe5b97e6f960..ee7140d8dda8 100644 --- a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_helpers.c +++ b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_helpers.c @@ -1378,12 +1378,13 @@ void dm_helpers_free_gpu_mem( bool dm_helpers_dmub_outbox_interrupt_control(struct dc_context *ctx, bool enable) { + struct amdgpu_device *adev = ctx->driver_context; enum dc_irq_source irq_source; bool ret; irq_source = DC_IRQ_SOURCE_DMCUB_OUTBOX; - ret = dc_interrupt_set(ctx->dc, irq_source, enable); + ret = amdgpu_dm_irq_set(adev, irq_source, enable); DRM_DEBUG_DRIVER("Dmub trace irq %sabling: r=%d\n", enable ? "en" : "dis", ret); diff --git a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.c b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.c index e49803a90eda..d59434e01dfd 100644 --- a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.c +++ b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.c @@ -424,6 +424,7 @@ int amdgpu_dm_irq_init(struct amdgpu_device *adev) DRM_DEBUG_KMS("DM_IRQ\n"); spin_lock_init(&adev->dm.irq_handler_list_table_lock); + spin_lock_init(&adev->dm.irq_reg_lock); for (src = 0; src < DAL_IRQ_SOURCES_NUMBER; src++) { /* low context handler list init */ @@ -496,7 +497,7 @@ void amdgpu_dm_irq_suspend(struct amdgpu_device *adev) hnd_list_l = &adev->dm.irq_handler_list_low_tab[src]; hnd_list_h = &adev->dm.irq_handler_list_high_tab[src]; if (!list_empty(hnd_list_l) || !list_empty(hnd_list_h)) - dc_interrupt_set(adev->dm.dc, src, false); + amdgpu_dm_irq_set(adev, src, false); DM_IRQ_TABLE_UNLOCK(adev, irq_table_flags); @@ -533,7 +534,7 @@ void amdgpu_dm_irq_resume_early(struct amdgpu_device *adev) hnd_list_l = &adev->dm.irq_handler_list_low_tab[src]; hnd_list_h = &adev->dm.irq_handler_list_high_tab[src]; if (!list_empty(hnd_list_l) || !list_empty(hnd_list_h)) - dc_interrupt_set(adev->dm.dc, src, true); + amdgpu_dm_irq_set(adev, src, true); } DM_IRQ_TABLE_UNLOCK(adev, irq_table_flags); @@ -558,7 +559,7 @@ void amdgpu_dm_irq_resume_late(struct amdgpu_device *adev) hnd_list_l = &adev->dm.irq_handler_list_low_tab[src]; hnd_list_h = &adev->dm.irq_handler_list_high_tab[src]; if (!list_empty(hnd_list_l) || !list_empty(hnd_list_h)) - dc_interrupt_set(adev->dm.dc, src, true); + amdgpu_dm_irq_set(adev, src, true); } DM_IRQ_TABLE_UNLOCK(adev, irq_table_flags); @@ -645,6 +646,21 @@ static void amdgpu_dm_irq_immediate_work(struct amdgpu_device *adev, DM_IRQ_TABLE_UNLOCK(adev, irq_table_flags); } +bool amdgpu_dm_irq_set(struct amdgpu_device *adev, enum dc_irq_source src, + bool enable) +{ + guard(spinlock_irqsave)(&adev->dm.irq_reg_lock); + + return dc_interrupt_set(adev->dm.dc, src, enable); +} + +void amdgpu_dm_irq_ack(struct amdgpu_device *adev, enum dc_irq_source src) +{ + guard(spinlock_irqsave)(&adev->dm.irq_reg_lock); + + dc_interrupt_ack(adev->dm.dc, src); +} + /** * amdgpu_dm_irq_handler - Generic DM IRQ handler * @adev: amdgpu base driver device containing the DM device @@ -665,7 +681,7 @@ static int amdgpu_dm_irq_handler(struct amdgpu_device *adev, entry->src_id, entry->src_data[0]); - dc_interrupt_ack(adev->dm.dc, src); + amdgpu_dm_irq_ack(adev, src); /* Call high irq work immediately */ amdgpu_dm_irq_immediate_work(adev, src); @@ -703,7 +719,7 @@ static int amdgpu_dm_set_hpd_irq_state(struct amdgpu_device *adev, enum dc_irq_source src = amdgpu_dm_hpd_to_dal_irq_source(type); bool st = (state == AMDGPU_IRQ_STATE_ENABLE); - dc_interrupt_set(adev->dm.dc, src, st); + amdgpu_dm_irq_set(adev, src, st); return 0; } @@ -737,7 +753,7 @@ static inline int dm_irq_state(struct amdgpu_device *adev, if (dc && dc->caps.ips_support && dc->idle_optimizations_allowed) dc_allow_idle_optimizations(dc, false); - dc_interrupt_set(adev->dm.dc, irq_source, st); + amdgpu_dm_irq_set(adev, irq_source, st); return 0; } @@ -791,7 +807,7 @@ static int amdgpu_dm_set_dmub_outbox_irq_state(struct amdgpu_device *adev, enum dc_irq_source irq_source = DC_IRQ_SOURCE_DMCUB_OUTBOX; bool st = (state == AMDGPU_IRQ_STATE_ENABLE); - dc_interrupt_set(adev->dm.dc, irq_source, st); + amdgpu_dm_irq_set(adev, irq_source, st); return 0; } @@ -817,7 +833,7 @@ static int amdgpu_dm_set_dmub_trace_irq_state(struct amdgpu_device *adev, enum dc_irq_source irq_source = DC_IRQ_SOURCE_DMCUB_OUTBOX0; bool st = (state == AMDGPU_IRQ_STATE_ENABLE); - dc_interrupt_set(adev->dm.dc, irq_source, st); + amdgpu_dm_irq_set(adev, irq_source, st); return 0; } @@ -881,9 +897,7 @@ void amdgpu_dm_set_irq_funcs(struct amdgpu_device *adev) } void amdgpu_dm_outbox_init(struct amdgpu_device *adev) { - dc_interrupt_set(adev->dm.dc, - DC_IRQ_SOURCE_DMCUB_OUTBOX, - true); + amdgpu_dm_irq_set(adev, DC_IRQ_SOURCE_DMCUB_OUTBOX, true); } /** @@ -905,7 +919,7 @@ void amdgpu_dm_hpd_init(struct amdgpu_device *adev) /* First, clear all hpd and hpdrx interrupts */ for (i = DC_IRQ_SOURCE_HPD1; i <= DC_IRQ_SOURCE_HPD6RX; i++) { - if (!dc_interrupt_set(adev->dm.dc, i, false)) + if (!amdgpu_dm_irq_set(adev, i, false)) drm_err(dev, "Failed to clear hpd(rx) source=%d on init\n", i); } @@ -934,7 +948,7 @@ void amdgpu_dm_hpd_init(struct amdgpu_device *adev) * of dm. Note that only hpd interrupt types are registered with * base driver; hpd_rx types aren't. IOW, amdgpu_irq_get/put on * hpd_rx isn't available. DM currently controls hpd_rx - * explicitly with dc_interrupt_set() + * explicitly with amdgpu_dm_irq_set() */ if (dc_link->irq_source_hpd != DC_IRQ_SOURCE_INVALID) { irq_type = dc_link->irq_source_hpd - DC_IRQ_SOURCE_HPD1; @@ -943,23 +957,21 @@ void amdgpu_dm_hpd_init(struct amdgpu_device *adev) * and what bios reports as the # of connectors with hpd * sources. Since the # of hpd source types registered * with base driver == mode_info.num_hpd, we have to - * fallback to dc_interrupt_set for the remaining types. + * fallback to amdgpu_dm_irq_set for the remaining types. */ if (irq_type < adev->mode_info.num_hpd) { if (amdgpu_irq_get(adev, &adev->hpd_irq, irq_type)) drm_err(dev, "DM_IRQ: Failed get HPD for source=%d)!\n", dc_link->irq_source_hpd); } else { - dc_interrupt_set(adev->dm.dc, - dc_link->irq_source_hpd, - true); + amdgpu_dm_irq_set(adev, dc_link->irq_source_hpd, + true); } } if (dc_link->irq_source_hpd_rx != DC_IRQ_SOURCE_INVALID) { - dc_interrupt_set(adev->dm.dc, - dc_link->irq_source_hpd_rx, - true); + amdgpu_dm_irq_set(adev, dc_link->irq_source_hpd_rx, + true); } } drm_connector_list_iter_end(&iter); @@ -1003,16 +1015,14 @@ void amdgpu_dm_hpd_fini(struct amdgpu_device *adev) drm_err(dev, "DM_IRQ: Failed put HPD for source=%d!\n", dc_link->irq_source_hpd); } else { - dc_interrupt_set(adev->dm.dc, - dc_link->irq_source_hpd, - false); + amdgpu_dm_irq_set(adev, dc_link->irq_source_hpd, + false); } } if (dc_link->irq_source_hpd_rx != DC_IRQ_SOURCE_INVALID) { - dc_interrupt_set(adev->dm.dc, - dc_link->irq_source_hpd_rx, - false); + amdgpu_dm_irq_set(adev, dc_link->irq_source_hpd_rx, + false); } } drm_connector_list_iter_end(&iter); diff --git a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.h b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.h index 4f6b58f4f90d..f0a4577848f3 100644 --- a/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.h +++ b/drivers/gpu/drm/amd/display/amdgpu_dm/amdgpu_dm_irq.h @@ -81,6 +81,34 @@ void amdgpu_dm_irq_unregister_interrupt(struct amdgpu_device *adev, enum dc_irq_source irq_source, void *ih_index); +/** + * amdgpu_dm_irq_set - enable or disable a DC interrupt source. + * + * @adev: AMD DRM device + * @src: DC interrupt source to toggle + * @enable: true to enable the source, false to disable it + * + * DM-wide replacement for dc_interrupt_set(). As locking is DM's + * responsibility, this is a thin wrapper serializes the underlying + * read-modify-write against the other interrupt sources sharing HW control + * registers with @src, so DM must never call dc_interrupt_set() directly. + * + * Returns: true if the source was toggled. + */ +bool amdgpu_dm_irq_set(struct amdgpu_device *adev, enum dc_irq_source src, + bool enable); + +/** + * amdgpu_dm_irq_ack - acknowledge a DC interrupt source. + * + * @adev: AMD DRM device + * @src: DC interrupt source to acknowledge + * + * DM-wide replacement for dc_interrupt_ack(), serialized the same way as + * amdgpu_dm_irq_set(). + */ +void amdgpu_dm_irq_ack(struct amdgpu_device *adev, enum dc_irq_source src); + void amdgpu_dm_set_irq_funcs(struct amdgpu_device *adev); void amdgpu_dm_outbox_init(struct amdgpu_device *adev); diff --git a/drivers/gpu/drm/amd/display/dc/dml/Makefile b/drivers/gpu/drm/amd/display/dc/dml/Makefile index 10d4ace04d4f..60bca8d14771 100644 --- a/drivers/gpu/drm/amd/display/dc/dml/Makefile +++ b/drivers/gpu/drm/amd/display/dc/dml/Makefile @@ -29,14 +29,18 @@ dml_ccflags := $(CC_FLAGS_FPU) dml_rcflags := $(CC_FLAGS_NO_FPU) ifneq ($(CONFIG_FRAME_WARN),0) - ifeq ($(filter y,$(CONFIG_KASAN)$(CONFIG_KCSAN)),y) + ifneq ($(filter y,$(CONFIG_KASAN) $(CONFIG_KCSAN) $(CONFIG_UBSAN)),) ifeq ($(CONFIG_CC_IS_CLANG)$(CONFIG_COMPILE_TEST),yy) frame_warn_limit := 4096 else frame_warn_limit := 3072 endif else - frame_warn_limit := 2048 + ifeq ($(CONFIG_CC_IS_CLANG),y) + frame_warn_limit := 3072 + else + frame_warn_limit := 2048 + endif endif ifeq ($(call test-lt, $(CONFIG_FRAME_WARN), $(frame_warn_limit)),y) diff --git a/drivers/gpu/drm/amd/display/dc/dml2_0/Makefile b/drivers/gpu/drm/amd/display/dc/dml2_0/Makefile index 44e00c2b7ac7..5f6bde11c3ca 100644 --- a/drivers/gpu/drm/amd/display/dc/dml2_0/Makefile +++ b/drivers/gpu/drm/amd/display/dc/dml2_0/Makefile @@ -28,7 +28,7 @@ dml2_ccflags := $(CC_FLAGS_FPU) dml2_rcflags := $(CC_FLAGS_NO_FPU) ifneq ($(CONFIG_FRAME_WARN),0) - ifeq ($(filter y,$(CONFIG_KASAN)$(CONFIG_KCSAN)),y) + ifneq ($(filter y,$(CONFIG_KASAN) $(CONFIG_KCSAN) $(CONFIG_UBSAN)),) ifeq ($(CONFIG_CC_IS_CLANG)$(CONFIG_COMPILE_TEST),yy) frame_warn_limit := 4096 else diff --git a/drivers/gpu/drm/amd/display/dc/link/link_dpms.c b/drivers/gpu/drm/amd/display/dc/link/link_dpms.c index ca5997748799..f49dba6b4e42 100644 --- a/drivers/gpu/drm/amd/display/dc/link/link_dpms.c +++ b/drivers/gpu/drm/amd/display/dc/link/link_dpms.c @@ -2296,6 +2296,8 @@ static enum dc_status enable_link( static bool allocate_usb4_bandwidth_for_stream(struct dc_stream_state *stream, int stream_bw) { + ASSERT(stream->sink); + struct dc_link *link = stream->sink->link; int req_bw = stream_bw; @@ -2342,6 +2344,8 @@ static bool allocate_usb4_bandwidth_for_stream(struct dc_stream_state *stream, i static bool allocate_usb4_bandwidth(struct dc_stream_state *stream) { + ASSERT(stream->sink); + bool ret; int bw = dc_bandwidth_in_kbps_from_timing(&stream->timing, @@ -2361,38 +2365,43 @@ static bool deallocate_usb4_bandwidth(struct dc_stream_state *stream) return ret; } +static struct vpg *get_vpg(struct pipe_ctx *pipe_ctx) +{ + if (dc_is_hdmi_frl_signal(pipe_ctx->stream->signal)) + return pipe_ctx->stream_res.hpo_frl_stream_enc->vpg; + else if (dp_is_128b_132b_signal(pipe_ctx)) + return pipe_ctx->stream_res.hpo_dp_stream_enc->vpg; + else + return pipe_ctx->stream_res.stream_enc->vpg; +} + void link_set_dpms_off(struct pipe_ctx *pipe_ctx) { - struct dc *dc = pipe_ctx->stream->ctx->dc; + DC_LOGGER_INIT(pipe_ctx->stream->ctx->logger); struct dc_stream_state *stream = pipe_ctx->stream; - struct dc_link *link = stream->sink->link; - struct vpg *vpg = pipe_ctx->stream_res.stream_enc->vpg; + struct dc *dc = stream->ctx->dc; + struct dc_link *link = stream->link; + struct vpg *vpg = get_vpg(pipe_ctx); enum dp_panel_mode panel_mode_dp = dp_get_panel_mode(link); - DC_LOGGER_INIT(pipe_ctx->stream->ctx->logger); - ASSERT(is_master_pipe_for_link(link, pipe_ctx)); - if (dp_is_128b_132b_signal(pipe_ctx)) - vpg = pipe_ctx->stream_res.hpo_dp_stream_enc->vpg; - if (dc_is_hdmi_frl_signal(pipe_ctx->stream->signal)) - vpg = pipe_ctx->stream_res.hpo_frl_stream_enc->vpg; - if (dc_is_virtual_signal(pipe_ctx->stream->signal)) + if (dc_is_virtual_signal(stream->signal)) return; - if (pipe_ctx->stream->sink) { - if (pipe_ctx->stream->sink->sink_signal != SIGNAL_TYPE_VIRTUAL && - pipe_ctx->stream->sink->sink_signal != SIGNAL_TYPE_NONE) { + if (stream->sink) { + if (stream->sink->sink_signal != SIGNAL_TYPE_VIRTUAL && + stream->sink->sink_signal != SIGNAL_TYPE_NONE) { DC_LOG_DC("%s pipe_ctx dispname=%s signal=%x link=%d sink_count=%d\n", __func__, - pipe_ctx->stream->sink->edid_caps.display_name, - pipe_ctx->stream->signal, link->link_index, link->sink_count); + stream->sink->edid_caps.display_name, + stream->signal, link->link_index, link->sink_count); } } link_wait_for_unlocked(link); - if (!pipe_ctx->stream->sink->edid_caps.panel_patch.skip_avmute) { - if (dc_is_hdmi_signal(pipe_ctx->stream->signal)) + if (stream->sink && !stream->sink->edid_caps.panel_patch.skip_avmute) { + if (dc_is_hdmi_signal(stream->signal)) set_avmute(pipe_ctx, true); } @@ -2402,15 +2411,15 @@ void link_set_dpms_off(struct pipe_ctx *pipe_ctx) dc->hwss.blank_stream(pipe_ctx); if (pipe_ctx->link_config.dp_tunnel_settings.should_use_dp_bw_allocation) - deallocate_usb4_bandwidth(pipe_ctx->stream); + deallocate_usb4_bandwidth(stream); - if (pipe_ctx->stream->signal == SIGNAL_TYPE_DISPLAY_PORT_MST) + if (stream->signal == SIGNAL_TYPE_DISPLAY_PORT_MST) deallocate_mst_payload(pipe_ctx); - else if (dc_is_dp_sst_signal(pipe_ctx->stream->signal) && + else if (dc_is_dp_sst_signal(stream->signal) && dp_is_128b_132b_signal(pipe_ctx)) update_sst_payload(pipe_ctx, false); - if (dc_is_hdmi_signal(pipe_ctx->stream->signal)) { + if (dc_is_hdmi_signal(stream->signal)) { struct ext_hdmi_settings settings = {0}; enum engine_id eng_id = pipe_ctx->stream_res.stream_enc->id; @@ -2433,7 +2442,7 @@ void link_set_dpms_off(struct pipe_ctx *pipe_ctx) } } - if (pipe_ctx->stream->signal == SIGNAL_TYPE_DISPLAY_PORT && + if (stream->signal == SIGNAL_TYPE_DISPLAY_PORT && !dp_is_128b_132b_signal(pipe_ctx)) { /* In DP1.x SST mode, our encoder will go to TPS1 @@ -2443,18 +2452,18 @@ void link_set_dpms_off(struct pipe_ctx *pipe_ctx) * state machine. * In DP2 or MST mode, our encoder will stay video active */ - disable_link(pipe_ctx->stream->link, &pipe_ctx->link_res, pipe_ctx->stream->signal); + disable_link(link, &pipe_ctx->link_res, stream->signal); dc->hwss.disable_stream(pipe_ctx); } else { dc->hwss.disable_stream(pipe_ctx); - disable_link(pipe_ctx->stream->link, &pipe_ctx->link_res, pipe_ctx->stream->signal); + disable_link(link, &pipe_ctx->link_res, stream->signal); } edp_set_panel_assr(link, pipe_ctx, &panel_mode_dp, false); - if (pipe_ctx->stream->timing.flags.DSC) { - if (dc_is_dp_signal(pipe_ctx->stream->signal)) + if (stream->timing.flags.DSC) { + if (dc_is_dp_signal(stream->signal)) link_set_dsc_enable(pipe_ctx, false); - else if (dc_is_hdmi_frl_signal(pipe_ctx->stream->signal)) + else if (dc_is_hdmi_frl_signal(stream->signal)) link_set_dsc_on_stream(pipe_ctx, false); } if (dp_is_128b_132b_signal(pipe_ctx)) { @@ -2468,7 +2477,7 @@ void link_set_dpms_off(struct pipe_ctx *pipe_ctx) /* for psp not exist case */ if (link->connector_signal == SIGNAL_TYPE_EDP && dc->debug.psp_disabled_wa) { /* reset internal save state to default since eDP is off */ - enum dp_panel_mode panel_mode = dp_get_panel_mode(pipe_ctx->stream->link); + enum dp_panel_mode panel_mode = dp_get_panel_mode(link); /* since current psp not loaded, we need to reset it to default */ link->panel_mode = panel_mode; } @@ -2478,62 +2487,57 @@ void link_set_dpms_on( struct dc_state *state, struct pipe_ctx *pipe_ctx) { - struct dc *dc = pipe_ctx->stream->ctx->dc; + DC_LOGGER_INIT(pipe_ctx->stream->ctx->logger); struct dc_stream_state *stream = pipe_ctx->stream; - struct dc_link *link = stream->sink->link; + struct dc *dc = stream->ctx->dc; + struct dc_link *link = stream->link; enum dc_status status; struct link_encoder *link_enc = pipe_ctx->link_res.dio_link_enc; enum otg_out_mux_dest otg_out_dest = OUT_MUX_DIO; - struct vpg *vpg = pipe_ctx->stream_res.stream_enc->vpg; + struct vpg *vpg = get_vpg(pipe_ctx); const struct link_hwss *link_hwss = get_link_hwss(link, &pipe_ctx->link_res); bool apply_edp_fast_boot_optimization = - pipe_ctx->stream->apply_edp_fast_boot_optimization; - - DC_LOGGER_INIT(pipe_ctx->stream->ctx->logger); + stream->apply_edp_fast_boot_optimization; ASSERT(is_master_pipe_for_link(link, pipe_ctx)); - if (dp_is_128b_132b_signal(pipe_ctx)) - vpg = pipe_ctx->stream_res.hpo_dp_stream_enc->vpg; - if (dc_is_hdmi_frl_signal(pipe_ctx->stream->signal)) - vpg = pipe_ctx->stream_res.hpo_frl_stream_enc->vpg; - if (dc_is_virtual_signal(pipe_ctx->stream->signal)) + if (dc_is_virtual_signal(stream->signal)) return; - if (pipe_ctx->stream->sink) { - if (pipe_ctx->stream->sink->sink_signal != SIGNAL_TYPE_VIRTUAL && - pipe_ctx->stream->sink->sink_signal != SIGNAL_TYPE_NONE) { + if (stream->sink) { + if (stream->sink->sink_signal != SIGNAL_TYPE_VIRTUAL && + stream->sink->sink_signal != SIGNAL_TYPE_NONE) { DC_LOG_DC("%s pipe_ctx dispname=%s signal=%x link=%d sink_count=%d\n", __func__, - pipe_ctx->stream->sink->edid_caps.display_name, - pipe_ctx->stream->signal, + stream->sink->edid_caps.display_name, + stream->signal, link->link_index, link->sink_count); } } - link_wait_for_unlocked(stream->link); + link_wait_for_unlocked(link); if (!dc->config.unify_link_enc_assignment) link_enc = link_enc_cfg_get_link_enc(link); ASSERT(link_enc); - if (!dc_is_virtual_signal(pipe_ctx->stream->signal) - && !dc_is_hdmi_frl_signal(pipe_ctx->stream->signal) + if (!dc_is_virtual_signal(stream->signal) + && !dc_is_hdmi_frl_signal(stream->signal) && !dp_is_128b_132b_signal(pipe_ctx)) { if (link_enc) link_enc->funcs->setup( link_enc, - pipe_ctx->stream->signal); + stream->signal); } - pipe_ctx->stream->link->link_state_valid = true; + link->link_state_valid = true; - if (dc_is_hdmi_frl_signal(pipe_ctx->stream->signal)) - hdmi_frl_decide_link_settings(stream, &stream->link->frl_link_settings, &pipe_ctx->dsc_padding_params); + if (dc_is_hdmi_frl_signal(stream->signal)) + hdmi_frl_decide_link_settings(stream, &link->frl_link_settings, &pipe_ctx->dsc_padding_params); if (pipe_ctx->stream_res.tg->funcs->set_out_mux) { if (dp_is_128b_132b_signal(pipe_ctx)) otg_out_dest = OUT_MUX_HPO_DP; - else if (dc_is_hdmi_frl_signal(pipe_ctx->stream->signal)) + else if (dc_is_hdmi_frl_signal(stream->signal)) otg_out_dest = OUT_MUX_HPO_FRL; else otg_out_dest = OUT_MUX_DIO; @@ -2542,7 +2546,7 @@ void link_set_dpms_on( link_hwss->setup_stream_attribute(pipe_ctx); - pipe_ctx->stream->apply_edp_fast_boot_optimization = false; + stream->apply_edp_fast_boot_optimization = false; // Enable VPG before building infoframe if (vpg && vpg->funcs->vpg_poweron) @@ -2551,15 +2555,15 @@ void link_set_dpms_on( resource_build_info_frame(pipe_ctx); dc->hwss.update_info_frame(pipe_ctx); - if (dc_is_dp_signal(pipe_ctx->stream->signal)) + if (dc_is_dp_signal(stream->signal)) dp_trace_source_sequence(link, DPCD_SOURCE_SEQ_AFTER_UPDATE_INFO_FRAME); /* Do not touch link on seamless boot optimization. */ - if (pipe_ctx->stream->apply_seamless_boot_optimization) { - pipe_ctx->stream->dpms_off = false; + if (stream->apply_seamless_boot_optimization) { + stream->dpms_off = false; /* Still enable stream features & audio on seamless boot for DP external displays */ - if (pipe_ctx->stream->signal == SIGNAL_TYPE_DISPLAY_PORT) { + if (stream->signal == SIGNAL_TYPE_DISPLAY_PORT) { enable_stream_features(pipe_ctx); dc->hwss.enable_audio_stream(pipe_ctx); } @@ -2569,11 +2573,11 @@ void link_set_dpms_on( } /* eDP lit up by bios already, no need to enable again. */ - if (pipe_ctx->stream->signal == SIGNAL_TYPE_EDP && + if (stream->signal == SIGNAL_TYPE_EDP && apply_edp_fast_boot_optimization && - !pipe_ctx->stream->timing.flags.DSC && + !stream->timing.flags.DSC && !pipe_ctx->next_odm_pipe) { - pipe_ctx->stream->dpms_off = false; + stream->dpms_off = false; update_psp_stream_config(pipe_ctx, false); if (link->is_dds) { @@ -2586,7 +2590,7 @@ void link_set_dpms_on( return; } - if (pipe_ctx->stream->dpms_off) + if (stream->dpms_off) return; /* For Dp tunneling link, a pending HPD means that we have a race condition between processing @@ -2604,9 +2608,9 @@ void link_set_dpms_on( * will be automatically set at a later time when the video is enabled * (DP_VID_STREAM_EN = 1). */ - if (pipe_ctx->stream->timing.flags.DSC) { - if (dc_is_dp_signal(pipe_ctx->stream->signal) || - dc_is_virtual_signal(pipe_ctx->stream->signal)) + if (stream->timing.flags.DSC) { + if (dc_is_dp_signal(stream->signal) || + dc_is_virtual_signal(stream->signal)) link_set_dsc_enable(pipe_ctx, true); } @@ -2617,7 +2621,7 @@ void link_set_dpms_on( if (status != DC_OK) { DC_LOG_WARNING("enabling link %u failed: %d\n", - pipe_ctx->stream->link->link_index, + link->link_index, status); /* Abort stream enable *unless* the failure was due to @@ -2627,19 +2631,18 @@ void link_set_dpms_on( */ if ((status != DC_FAIL_DP_LINK_TRAINING && status != DC_FAIL_HDMI_FRL_LINK_TRAINING) || - pipe_ctx->stream->signal == SIGNAL_TYPE_DISPLAY_PORT_MST) { - if (false == stream->link->link_status.link_active) - disable_link(stream->link, &pipe_ctx->link_res, - pipe_ctx->stream->signal); + stream->signal == SIGNAL_TYPE_DISPLAY_PORT_MST) { + if (false == link->link_status.link_active) + disable_link(link, &pipe_ctx->link_res, + stream->signal); BREAK_TO_DEBUGGER(); return; } } - if (pipe_ctx->stream->timing.flags.DSC && - dc_is_hdmi_frl_signal(pipe_ctx->stream->signal)) - //TODO: bring HDMI FRL in line with DP - link_set_dsc_on_stream(pipe_ctx, true); + if (stream->timing.flags.DSC && dc_is_hdmi_frl_signal(stream->signal)) + //TODO: bring HDMI FRL in line with DP + link_set_dsc_on_stream(pipe_ctx, true); /* turn off otg test pattern if enable */ if (pipe_ctx->stream_res.tg->funcs->set_test_pattern) @@ -2651,37 +2654,37 @@ void link_set_dpms_on( * as a workaround for the incorrect value being applied * from transmitter control. */ - if (!(dc_is_virtual_signal(pipe_ctx->stream->signal) || - dc_is_hdmi_frl_signal(pipe_ctx->stream->signal) || + if (!(dc_is_virtual_signal(stream->signal) || + dc_is_hdmi_frl_signal(stream->signal) || dp_is_128b_132b_signal(pipe_ctx))) { if (link_enc) link_enc->funcs->setup( link_enc, - pipe_ctx->stream->signal); + stream->signal); } dc->hwss.enable_stream(pipe_ctx); /* Set DPS PPS SDP (AKA "info frames") */ - if (pipe_ctx->stream->timing.flags.DSC) { - if (dc_is_dp_signal(pipe_ctx->stream->signal) || - dc_is_virtual_signal(pipe_ctx->stream->signal)) { + if (stream->timing.flags.DSC) { + if (dc_is_dp_signal(stream->signal) || + dc_is_virtual_signal(stream->signal)) { dp_set_dsc_on_rx(pipe_ctx, true); link_set_dsc_pps_packet(pipe_ctx, true, true); } } - if (dc_is_dp_signal(pipe_ctx->stream->signal)) + if (dc_is_dp_signal(stream->signal)) dp_set_hblank_reduction_on_rx(pipe_ctx); if (pipe_ctx->link_config.dp_tunnel_settings.should_use_dp_bw_allocation) - allocate_usb4_bandwidth(pipe_ctx->stream); + allocate_usb4_bandwidth(stream); - if (pipe_ctx->stream->signal == SIGNAL_TYPE_DISPLAY_PORT_MST) + if (stream->signal == SIGNAL_TYPE_DISPLAY_PORT_MST) allocate_mst_payload(pipe_ctx); - else if (dc_is_dp_sst_signal(pipe_ctx->stream->signal) && + else if (dc_is_dp_sst_signal(stream->signal) && dp_is_128b_132b_signal(pipe_ctx)) update_sst_payload(pipe_ctx, true); @@ -2690,23 +2693,22 @@ void link_set_dpms_on( * training and stream unblank resolves the corruption issue. * This is workaround. */ - if (pipe_ctx->stream->signal == SIGNAL_TYPE_EDP && + if (stream->signal == SIGNAL_TYPE_EDP && link->is_display_mux_present) msleep(20); dc->hwss.unblank_stream(pipe_ctx, - &pipe_ctx->stream->link->cur_link_settings); + &link->cur_link_settings); if (stream->sink_patches.delay_ignore_msa > 0) msleep(stream->sink_patches.delay_ignore_msa); - if (dc_is_dp_signal(pipe_ctx->stream->signal)) + if (dc_is_dp_signal(stream->signal)) enable_stream_features(pipe_ctx); update_psp_stream_config(pipe_ctx, false); dc->hwss.enable_audio_stream(pipe_ctx); - if (dc_is_hdmi_signal(pipe_ctx->stream->signal)) { + if (dc_is_hdmi_signal(stream->signal)) set_avmute(pipe_ctx, false); - } } diff --git a/drivers/gpu/drm/bridge/samsung-dsim.c b/drivers/gpu/drm/bridge/samsung-dsim.c index 9ee0515074c7..7e320eac2e4a 100644 --- a/drivers/gpu/drm/bridge/samsung-dsim.c +++ b/drivers/gpu/drm/bridge/samsung-dsim.c @@ -1862,7 +1862,7 @@ static int samsung_dsim_register_te_irq(struct samsung_dsim *dsi, struct device int te_gpio_irq; int ret; - dsi->te_gpio = devm_gpiod_get_optional(dev, "te", GPIOD_IN); + dsi->te_gpio = gpiod_get_optional(dev, "te", GPIOD_IN); if (!dsi->te_gpio) return 0; else if (IS_ERR(dsi->te_gpio)) diff --git a/drivers/gpu/drm/clients/drm_fbdev_client.c b/drivers/gpu/drm/clients/drm_fbdev_client.c index 91d196a397cf..1c16bc1084c4 100644 --- a/drivers/gpu/drm/clients/drm_fbdev_client.c +++ b/drivers/gpu/drm/clients/drm_fbdev_client.c @@ -42,6 +42,14 @@ static int drm_fbdev_client_restore(struct drm_client_dev *client, bool force) { struct drm_fb_helper *fb_helper = drm_fb_helper_from_client(client); + /* + * The client is registered before the initial fbdev probe. + * If probing failed, the client remains registered but there + * is no valid fbdev framebuffer to restore. + */ + if (!fb_helper->info || !fb_helper->fb) + return 0; + drm_fb_helper_restore_fbdev_mode_unlocked(fb_helper, force); return 0; diff --git a/drivers/gpu/drm/drm_pagemap.c b/drivers/gpu/drm/drm_pagemap.c index 89e1ddff84dd..a0546955d0b9 100644 --- a/drivers/gpu/drm/drm_pagemap.c +++ b/drivers/gpu/drm/drm_pagemap.c @@ -1368,13 +1368,13 @@ int drm_pagemap_evict_to_ram(struct drm_pagemap_devmem *devmem_allocation) goto err_finalize; err_finalize: + drm_pagemap_migrate_unmap_pages(devmem_allocation->dev, pagemap_addr, dst, npages, + DMA_FROM_DEVICE, &state); if (err) drm_pagemap_migration_unlock_put_pages(npages, dst); migrate_device_pages(src, dst, npages); drm_pagemap_retire_migrated_pages(src, npages); migrate_device_finalize(src, dst, npages); - drm_pagemap_migrate_unmap_pages(devmem_allocation->dev, pagemap_addr, dst, npages, - DMA_FROM_DEVICE, &state); err_free: kvfree(buf); @@ -1494,15 +1494,15 @@ static int __drm_pagemap_migrate_to_ram(struct vm_area_struct *vas, goto err_finalize; err_finalize: + if (dev) + drm_pagemap_migrate_unmap_pages(dev, pagemap_addr, migrate.dst, + npages, DMA_FROM_DEVICE, + &state); if (err) drm_pagemap_migration_unlock_put_pages(npages, migrate.dst); migrate_vma_pages(&migrate); drm_pagemap_retire_migrated_pages(migrate.src, npages); migrate_vma_finalize(&migrate); - if (dev) - drm_pagemap_migrate_unmap_pages(dev, pagemap_addr, migrate.dst, - npages, DMA_FROM_DEVICE, - &state); err_free: kvfree(buf); err_out: diff --git a/drivers/gpu/drm/i915/display/intel_cursor.c b/drivers/gpu/drm/i915/display/intel_cursor.c index 52347668f27d..dad0e0eb0e0c 100644 --- a/drivers/gpu/drm/i915/display/intel_cursor.c +++ b/drivers/gpu/drm/i915/display/intel_cursor.c @@ -536,7 +536,8 @@ static void i9xx_cursor_disable_sel_fetch_arm(struct intel_dsb *dsb, struct intel_display *display = to_intel_display(plane); enum pipe pipe = plane->pipe; - if (!crtc_state->enable_psr2_sel_fetch) + if (!crtc_state->enable_psr2_sel_fetch && + !crtc_state->clear_psr2_sel_fetch) return; intel_de_write_dsb(display, dsb, SEL_FETCH_CUR_CTL(pipe), 0); @@ -569,8 +570,10 @@ static void i9xx_cursor_update_sel_fetch_arm(struct intel_dsb *dsb, struct intel_display *display = to_intel_display(plane); enum pipe pipe = plane->pipe; - if (!crtc_state->enable_psr2_sel_fetch) + if (!crtc_state->enable_psr2_sel_fetch) { + i9xx_cursor_disable_sel_fetch_arm(dsb, plane, crtc_state); return; + } if (drm_rect_height(&plane_state->psr2_sel_fetch_area) > 0) { if (crtc_state->enable_psr2_su_region_et) { diff --git a/drivers/gpu/drm/i915/display/intel_display_types.h b/drivers/gpu/drm/i915/display/intel_display_types.h index 96422641ae44..18e6b01efd38 100644 --- a/drivers/gpu/drm/i915/display/intel_display_types.h +++ b/drivers/gpu/drm/i915/display/intel_display_types.h @@ -1177,6 +1177,8 @@ struct intel_crtc_state { bool has_sel_update; bool enable_psr2_sel_fetch; bool enable_psr2_su_region_et; + /* Drop the stale selective fetch enable bits as selective fetch is turned off */ + bool clear_psr2_sel_fetch; bool req_psr2_sdp_prior_scanline; bool has_panel_replay; bool link_off_after_as_sdp_when_pr_active; diff --git a/drivers/gpu/drm/i915/display/intel_dp_mst.c b/drivers/gpu/drm/i915/display/intel_dp_mst.c index bcdc50491347..a1922f95e5d5 100644 --- a/drivers/gpu/drm/i915/display/intel_dp_mst.c +++ b/drivers/gpu/drm/i915/display/intel_dp_mst.c @@ -813,7 +813,8 @@ static u8 get_pipes_downstream_of_mst_port(struct intel_atomic_state *state, if (&connector->mst.dp->mst.mgr != mst_mgr) continue; - if (connector->mst.port != parent_port && + if (parent_port && + connector->mst.port != parent_port && !drm_dp_mst_port_downstream_of_parent(mst_mgr, connector->mst.port, parent_port)) @@ -2121,6 +2122,27 @@ bool intel_dp_mst_crtc_needs_modeset(struct intel_atomic_state *state, return false; } +bool intel_dp_mst_stream_disconnected(struct intel_atomic_state *state, + const struct intel_crtc *crtc) +{ + struct intel_connector *connector; + + connector = get_connector_in_state_for_crtc(state, crtc); + if (!connector) + return false; + + if (!connector->mst.dp) + return false; + + if (!connector->mst.dp->mst.mgr.mst_state) + return true; + + if (drm_connector_is_unregistered(&connector->base)) + return true; + + return false; +} + /** * intel_dp_mst_prepare_probe - Prepare an MST link for topology probing * @intel_dp: DP port object diff --git a/drivers/gpu/drm/i915/display/intel_dp_mst.h b/drivers/gpu/drm/i915/display/intel_dp_mst.h index ab09b487c6bb..8ce89242c05c 100644 --- a/drivers/gpu/drm/i915/display/intel_dp_mst.h +++ b/drivers/gpu/drm/i915/display/intel_dp_mst.h @@ -28,6 +28,8 @@ int intel_dp_mst_atomic_check_link(struct intel_atomic_state *state, struct intel_link_bw_limits *limits); bool intel_dp_mst_crtc_needs_modeset(struct intel_atomic_state *state, struct intel_crtc *crtc); +bool intel_dp_mst_stream_disconnected(struct intel_atomic_state *state, + const struct intel_crtc *crtc); void intel_dp_mst_prepare_probe(struct intel_dp *intel_dp); bool intel_dp_mst_verify_dpcd_state(struct intel_dp *intel_dp); diff --git a/drivers/gpu/drm/i915/display/intel_link_bw.c b/drivers/gpu/drm/i915/display/intel_link_bw.c index b47474a3e9fe..e71e76d6fd3e 100644 --- a/drivers/gpu/drm/i915/display/intel_link_bw.c +++ b/drivers/gpu/drm/i915/display/intel_link_bw.c @@ -64,7 +64,8 @@ void intel_link_bw_init_limits(struct intel_atomic_state *state, intel_atomic_get_new_crtc_state(state, crtc); int forced_bpp_x16 = get_forced_link_bpp_x16(state, crtc); - if (state->base.duplicated && crtc_state) { + if ((state->base.duplicated && crtc_state) || + intel_dp_mst_stream_disconnected(state, crtc)) { limits->max_bpp_x16[pipe] = crtc_state->max_link_bpp_x16; if (intel_dsc_enabled_on_link(crtc_state)) limits->link_dsc_pipes |= BIT(pipe); diff --git a/drivers/gpu/drm/i915/display/intel_psr.c b/drivers/gpu/drm/i915/display/intel_psr.c index beaa1d62613d..87128104476a 100644 --- a/drivers/gpu/drm/i915/display/intel_psr.c +++ b/drivers/gpu/drm/i915/display/intel_psr.c @@ -2933,6 +2933,8 @@ int intel_psr2_sel_fetch_update(struct intel_atomic_state *state, struct intel_crtc *crtc) { struct intel_display *display = to_intel_display(state); + const struct intel_crtc_state *old_crtc_state = + intel_atomic_get_old_crtc_state(state, crtc); struct intel_crtc_state *crtc_state = intel_atomic_get_new_crtc_state(state, crtc); struct intel_plane_state *new_plane_state, *old_plane_state; struct intel_plane *plane; @@ -2945,6 +2947,19 @@ int intel_psr2_sel_fetch_update(struct intel_atomic_state *state, bool full_update = false, su_area_changed; int i, ret; + /* + * Selective fetch is not always usable, for instance it is dropped + * while pipe CRC is active. The planes keep their selective fetch + * enable bit set in hardware over that, and a plane disabled while + * selective fetch is off never gets the bit cleared. Once selective + * fetch comes back the hardware would resume fetching for a plane that + * is no longer enabled and keep its DDB range reserved, so have the + * plane update drop the bit for every plane of the pipe as selective + * fetch is turned off. + */ + crtc_state->clear_psr2_sel_fetch = old_crtc_state->enable_psr2_sel_fetch && + !crtc_state->enable_psr2_sel_fetch; + if (!crtc_state->enable_psr2_sel_fetch) return 0; diff --git a/drivers/gpu/drm/i915/display/intel_quirks.c b/drivers/gpu/drm/i915/display/intel_quirks.c index 33245f44c0d5..7d7db774d8c7 100644 --- a/drivers/gpu/drm/i915/display/intel_quirks.c +++ b/drivers/gpu/drm/i915/display/intel_quirks.c @@ -257,6 +257,9 @@ static struct intel_quirk intel_quirks[] = { /* Dell XPS 13 7390 2-in-1 */ { 0x8a52, 0x1028, 0x08b0, quirk_edp_limit_rate_hbr2 }, + /* HP Pavilion Plus Laptop 14-ew1xxx */ + { 0x7d55, 0x103c, 0x8c31, quirk_edp_limit_rate_hbr2 }, + /* Xiaomi Book Pro 14 2026 */ { 0xb081, 0x1d72, 0x2424, quirk_disable_psr2 }, }; diff --git a/drivers/gpu/drm/i915/display/skl_universal_plane.c b/drivers/gpu/drm/i915/display/skl_universal_plane.c index 164b7d61c9a3..3dad2da4c3aa 100644 --- a/drivers/gpu/drm/i915/display/skl_universal_plane.c +++ b/drivers/gpu/drm/i915/display/skl_universal_plane.c @@ -885,7 +885,8 @@ static void icl_plane_disable_sel_fetch_arm(struct intel_dsb *dsb, struct intel_display *display = to_intel_display(plane); enum pipe pipe = plane->pipe; - if (!crtc_state->enable_psr2_sel_fetch) + if (!crtc_state->enable_psr2_sel_fetch && + !crtc_state->clear_psr2_sel_fetch) return; intel_de_write_dsb(display, dsb, SEL_FETCH_PLANE_CTL(pipe, plane->id), 0); @@ -1634,10 +1635,8 @@ static void icl_plane_update_sel_fetch_arm(struct intel_dsb *dsb, struct intel_display *display = to_intel_display(plane); enum pipe pipe = plane->pipe; - if (!crtc_state->enable_psr2_sel_fetch) - return; - - if (drm_rect_height(&plane_state->psr2_sel_fetch_area) > 0) + if (crtc_state->enable_psr2_sel_fetch && + drm_rect_height(&plane_state->psr2_sel_fetch_area) > 0) intel_de_write_dsb(display, dsb, SEL_FETCH_PLANE_CTL(pipe, plane->id), SEL_FETCH_PLANE_CTL_ENABLE); else diff --git a/drivers/gpu/drm/i915/gem/i915_gem_object.c b/drivers/gpu/drm/i915/gem/i915_gem_object.c index 5172d3982654..9e01f8b2079a 100644 --- a/drivers/gpu/drm/i915/gem/i915_gem_object.c +++ b/drivers/gpu/drm/i915/gem/i915_gem_object.c @@ -89,6 +89,7 @@ struct drm_i915_gem_object *i915_gem_object_alloc(void) void i915_gem_object_free(struct drm_i915_gem_object *obj) { + dma_resv_fini(&obj->base._resv); return kmem_cache_free(slab_objects, obj); } @@ -144,7 +145,6 @@ void __i915_gem_object_fini(struct drm_i915_gem_object *obj) { mutex_destroy(&obj->mm.get_page.lock); mutex_destroy(&obj->mm.get_dma_page.lock); - dma_resv_fini(&obj->base._resv); } /** diff --git a/drivers/gpu/drm/imagination/pvr_free_list.c b/drivers/gpu/drm/imagination/pvr_free_list.c index e85cac83834c..faf5e586d8dc 100644 --- a/drivers/gpu/drm/imagination/pvr_free_list.c +++ b/drivers/gpu/drm/imagination/pvr_free_list.c @@ -8,6 +8,7 @@ #include "pvr_vm.h" #include +#include #include #include #include @@ -612,13 +613,21 @@ pvr_free_list_process_reconstruct_req(struct pvr_device *pvr_dev, }; struct rogue_fwif_freelists_reconstruction_data *resp = &resp_cmd.cmd_data.free_lists_reconstruction_data; + u32 count = min_t(u32, req->freelist_count, + ARRAY_SIZE(req->freelist_ids)); - for (u32 i = 0; i < req->freelist_count; i++) + if (count != req->freelist_count) { + drm_warn_once(from_pvr_device(pvr_dev), + "Requested reconstruction of %u freelists, limiting to %u\n", + req->freelist_count, count); + } + + for (u32 i = 0; i < count; i++) pvr_free_list_reconstruct(pvr_dev, req->freelist_ids[i]); - resp->freelist_count = req->freelist_count; + resp->freelist_count = count; memcpy(resp->freelist_ids, req->freelist_ids, - req->freelist_count * sizeof(resp->freelist_ids[0])); + count * sizeof(resp->freelist_ids[0])); WARN_ON(pvr_kccb_send_cmd(pvr_dev, &resp_cmd, NULL)); } diff --git a/drivers/gpu/drm/imagination/pvr_mmu.c b/drivers/gpu/drm/imagination/pvr_mmu.c index 3cac482e1034..62eae7fcd5a2 100644 --- a/drivers/gpu/drm/imagination/pvr_mmu.c +++ b/drivers/gpu/drm/imagination/pvr_mmu.c @@ -12,6 +12,7 @@ #include "pvr_rogue_mmu_defs.h" #include +#include #include #include #include @@ -2335,6 +2336,7 @@ void pvr_mmu_op_context_destroy(struct pvr_mmu_op_context *op_ctx) * pvr_mmu_op_context_create() - Create an MMU op context. * @ctx: MMU context associated with owning VM context. * @sgt: Scatter gather table containing pages pinned for use by this context. + * @device_addr: Virtual device address at the start of the requested mapping. * @sgt_offset: Start offset of the requested device-virtual memory mapping. * @size: Size in bytes of the requested device-virtual memory mapping. For an * unmapping, this should be zero so that no page tables are allocated. @@ -2346,8 +2348,9 @@ void pvr_mmu_op_context_destroy(struct pvr_mmu_op_context *op_ctx) */ struct pvr_mmu_op_context * pvr_mmu_op_context_create(struct pvr_mmu_context *ctx, struct sg_table *sgt, - u64 sgt_offset, u64 size) + u64 device_addr, u64 sgt_offset, u64 size) { + u64 start_addr = device_addr + sgt_offset; int err; struct pvr_mmu_op_context *op_ctx = kzalloc_obj(*op_ctx); @@ -2363,16 +2366,16 @@ pvr_mmu_op_context_create(struct pvr_mmu_context *ctx, struct sg_table *sgt, if (size) { /* * The number of page table objects we need to prealloc is - * indicated by the mapping size, start offset and the sizes + * indicated by the mapping size, start address and the sizes * of the areas mapped per PT or PD. The range calculation is * identical to that for the index into a table for a device * address, so we reuse those functions here. */ - const u32 l1_start_idx = pvr_page_table_l2_idx(sgt_offset); - const u32 l1_end_idx = pvr_page_table_l2_idx(sgt_offset + size); + const u32 l1_start_idx = pvr_page_table_l2_idx(start_addr); + const u32 l1_end_idx = pvr_page_table_l2_idx(start_addr + size); const u32 l1_count = l1_end_idx - l1_start_idx + 1; - const u32 l0_start_idx = pvr_page_table_l1_idx(sgt_offset); - const u32 l0_end_idx = pvr_page_table_l1_idx(sgt_offset + size); + const u32 l0_start_idx = pvr_page_table_l1_idx(start_addr); + const u32 l0_end_idx = pvr_page_table_l1_idx(start_addr + size); const u32 l0_count = l0_end_idx - l0_start_idx + 1; /* @@ -2553,7 +2556,9 @@ pvr_mmu_map_sgl(struct pvr_mmu_op_context *op_ctx, struct scatterlist *sgl, err_destroy_pages: memcpy(&op_ctx->curr_page, &ptr_copy, sizeof(op_ctx->curr_page)); - err = pvr_mmu_op_context_unmap_curr_page(op_ctx, page); + if (pvr_mmu_op_context_unmap_curr_page(op_ctx, page)) + drm_err(from_pvr_device(op_ctx->mmu_ctx->pvr_dev), + "%s : Failure in unmapping pages\n", __func__); return err; } diff --git a/drivers/gpu/drm/imagination/pvr_mmu.h b/drivers/gpu/drm/imagination/pvr_mmu.h index a8ecd460168d..2c02d61ba0a2 100644 --- a/drivers/gpu/drm/imagination/pvr_mmu.h +++ b/drivers/gpu/drm/imagination/pvr_mmu.h @@ -99,7 +99,7 @@ dma_addr_t pvr_mmu_get_root_table_dma_addr(struct pvr_mmu_context *ctx); void pvr_mmu_op_context_destroy(struct pvr_mmu_op_context *op_ctx); struct pvr_mmu_op_context * pvr_mmu_op_context_create(struct pvr_mmu_context *ctx, - struct sg_table *sgt, u64 sgt_offset, u64 size); + struct sg_table *sgt, u64 device_addr, u64 sgt_offset, u64 size); int pvr_mmu_map(struct pvr_mmu_op_context *op_ctx, u64 size, u64 flags, u64 device_addr); diff --git a/drivers/gpu/drm/imagination/pvr_vm.c b/drivers/gpu/drm/imagination/pvr_vm.c index ceb78694cd98..55cc999f3708 100644 --- a/drivers/gpu/drm/imagination/pvr_vm.c +++ b/drivers/gpu/drm/imagination/pvr_vm.c @@ -276,7 +276,7 @@ pvr_vm_bind_op_map_init(struct pvr_vm_bind_op *bind_op, goto err_bind_op_fini; bind_op->mmu_op_ctx = - pvr_mmu_op_context_create(vm_ctx->mmu_ctx, sgt, offset, size); + pvr_mmu_op_context_create(vm_ctx->mmu_ctx, sgt, device_addr, offset, size); err = PTR_ERR_OR_ZERO(bind_op->mmu_op_ctx); if (err) { bind_op->mmu_op_ctx = NULL; @@ -318,7 +318,7 @@ pvr_vm_bind_op_unmap_init(struct pvr_vm_bind_op *bind_op, } bind_op->mmu_op_ctx = - pvr_mmu_op_context_create(vm_ctx->mmu_ctx, NULL, 0, 0); + pvr_mmu_op_context_create(vm_ctx->mmu_ctx, NULL, device_addr, 0, 0); err = PTR_ERR_OR_ZERO(bind_op->mmu_op_ctx); if (err) { bind_op->mmu_op_ctx = NULL; diff --git a/drivers/gpu/drm/nouveau/nouveau_bo.c b/drivers/gpu/drm/nouveau/nouveau_bo.c index 0e8de6d4b36f..6dcb92575eb4 100644 --- a/drivers/gpu/drm/nouveau/nouveau_bo.c +++ b/drivers/gpu/drm/nouveau/nouveau_bo.c @@ -578,8 +578,9 @@ int nouveau_bo_pin_locked(struct nouveau_bo *nvbo, uint32_t domain, bool contig) "0x%08x vs 0x%08x\n", bo, bo->resource->mem_type, domain); ret = -EBUSY; + } else { + ttm_bo_pin(&nvbo->bo); } - ttm_bo_pin(&nvbo->bo); goto out; } diff --git a/drivers/gpu/drm/nouveau/nouveau_connector.c b/drivers/gpu/drm/nouveau/nouveau_connector.c index b0b0ad9a0c24..4cfc9c7c2ae0 100644 --- a/drivers/gpu/drm/nouveau/nouveau_connector.c +++ b/drivers/gpu/drm/nouveau/nouveau_connector.c @@ -600,8 +600,11 @@ nouveau_connector_detect(struct drm_connector *connector, bool force) new_edid = drm_get_edid(connector, nv_encoder->i2c); } else { ret = nvif_outp_edid_get(&nv_encoder->outp, (u8 **)&new_edid); - if (ret < 0) + if (ret < 0) { + pm_runtime_mark_last_busy(dev->dev); + pm_runtime_put_autosuspend(dev->dev); return connector_status_disconnected; + } } nouveau_connector_set_edid(nv_connector, new_edid); diff --git a/drivers/gpu/drm/nouveau/nouveau_dmem.c b/drivers/gpu/drm/nouveau/nouveau_dmem.c index ad4570c50be7..e74d7bb975a8 100644 --- a/drivers/gpu/drm/nouveau/nouveau_dmem.c +++ b/drivers/gpu/drm/nouveau/nouveau_dmem.c @@ -339,8 +339,8 @@ nouveau_dmem_chunk_alloc(struct nouveau_drm *drm, struct page **ppage, chunk->pagemap.ops = &nouveau_dmem_pagemap_ops; chunk->pagemap.owner = drm->dev; - ret = nouveau_bo_new_pin(&drm->client, NOUVEAU_GEM_DOMAIN_VRAM, DMEM_CHUNK_SIZE, - &chunk->bo); + ret = nouveau_bo_new_pin(&drm->client, NOUVEAU_GEM_DOMAIN_VRAM, + DMEM_CHUNK_SIZE * NR_CHUNKS, &chunk->bo); if (ret) goto out_release; diff --git a/drivers/gpu/drm/nouveau/nouveau_drm.c b/drivers/gpu/drm/nouveau/nouveau_drm.c index e16f59b00f6f..c03638ebd352 100644 --- a/drivers/gpu/drm/nouveau/nouveau_drm.c +++ b/drivers/gpu/drm/nouveau/nouveau_drm.c @@ -585,6 +585,7 @@ nouveau_drm_device_fini(struct nouveau_drm *drm) if (nouveau_pmops_runtime()) { pm_runtime_get_sync(dev->dev); pm_runtime_forbid(dev->dev); + pm_runtime_dont_use_autosuspend(dev->dev); } nouveau_led_fini(dev); @@ -1251,10 +1252,8 @@ nouveau_drm_open(struct drm_device *dev, struct drm_file *fpriv) mutex_unlock(&drm->clients_lock); done: - if (ret && cli) { - nouveau_cli_fini(cli); + if (ret && cli) kfree(cli); - } pm_runtime_mark_last_busy(dev->dev); pm_runtime_put_autosuspend(dev->dev); diff --git a/drivers/gpu/drm/nouveau/nouveau_gem.c b/drivers/gpu/drm/nouveau/nouveau_gem.c index c5a24dff4b69..1958bd52dbbc 100644 --- a/drivers/gpu/drm/nouveau/nouveau_gem.c +++ b/drivers/gpu/drm/nouveau/nouveau_gem.c @@ -522,6 +522,7 @@ validate_init(struct nouveau_channel *chan, struct drm_file *file_priv, if (unlikely(ret)) { if (ret != -ERESTARTSYS) NV_PRINTK(err, cli, "fail reserve\n"); + drm_gem_object_put(gem); break; } } @@ -531,6 +532,7 @@ validate_init(struct nouveau_channel *chan, struct drm_file *file_priv, struct nouveau_vma *vma = nouveau_vma_find(nvbo, vmm); if (!vma) { NV_PRINTK(err, cli, "vma not found!\n"); + drm_gem_object_put(gem); ret = -EINVAL; break; } diff --git a/drivers/gpu/drm/nouveau/nouveau_sched.c b/drivers/gpu/drm/nouveau/nouveau_sched.c index 8b9f935afe09..b3f02c490ecb 100644 --- a/drivers/gpu/drm/nouveau/nouveau_sched.c +++ b/drivers/gpu/drm/nouveau/nouveau_sched.c @@ -517,7 +517,7 @@ nouveau_sched_destroy(struct nouveau_sched **psched) struct nouveau_sched *sched = *psched; nouveau_sched_fini(sched); - kfree(sched); + kfree_rcu(sched, rcu); *psched = NULL; } diff --git a/drivers/gpu/drm/nouveau/nouveau_sched.h b/drivers/gpu/drm/nouveau/nouveau_sched.h index 20cd1da8db73..51ce8dcf6285 100644 --- a/drivers/gpu/drm/nouveau/nouveau_sched.h +++ b/drivers/gpu/drm/nouveau/nouveau_sched.h @@ -98,6 +98,7 @@ void nouveau_job_free(struct nouveau_job *job); struct nouveau_sched { struct drm_gpu_scheduler base; + struct rcu_head rcu; struct drm_sched_entity entity; struct workqueue_struct *wq; struct mutex mutex; diff --git a/drivers/gpu/drm/nouveau/nouveau_uvmm.c b/drivers/gpu/drm/nouveau/nouveau_uvmm.c index fc125fd44a9b..2026fe6b48c6 100644 --- a/drivers/gpu/drm/nouveau/nouveau_uvmm.c +++ b/drivers/gpu/drm/nouveau/nouveau_uvmm.c @@ -846,6 +846,9 @@ op_map(struct nouveau_uvma *uvma) { struct nouveau_bo *nvbo = nouveau_gem_object(uvma->va.gem.obj); + if (drm_gpuva_invalidated(&uvma->va)) + return; + nouveau_uvma_map(uvma, nouveau_mem(nvbo->bo.resource)); } @@ -1232,6 +1235,7 @@ bind_lock_validate(struct nouveau_job *job, struct drm_exec *exec, drm_gpuva_for_each_op(va_op, op->ops) { struct drm_gem_object *obj = op_gem_obj(va_op); + struct nouveau_bo *nvbo; if (unlikely(!obj)) continue; @@ -1246,8 +1250,13 @@ bind_lock_validate(struct nouveau_job *job, struct drm_exec *exec, if (va_op->op == DRM_GPUVA_OP_UNMAP) continue; - ret = nouveau_bo_validate(nouveau_gem_object(obj), - true, false); + nvbo = nouveau_gem_object(obj); + if (!(nvbo->valid_domains & + (NOUVEAU_GEM_DOMAIN_VRAM | NOUVEAU_GEM_DOMAIN_GART))) + return -EINVAL; + + nouveau_bo_placement_set(nvbo, nvbo->valid_domains, 0); + ret = nouveau_bo_validate(nvbo, true, false); if (ret) return ret; } diff --git a/drivers/gpu/drm/nouveau/nvif/vmm.c b/drivers/gpu/drm/nouveau/nvif/vmm.c index 65c3e883b119..579af70766f2 100644 --- a/drivers/gpu/drm/nouveau/nvif/vmm.c +++ b/drivers/gpu/drm/nouveau/nvif/vmm.c @@ -192,6 +192,7 @@ void nvif_vmm_dtor(struct nvif_vmm *vmm) { kfree(vmm->page); + vmm->page = NULL; nvif_object_dtor(&vmm->object); } diff --git a/drivers/gpu/drm/nouveau/nvkm/engine/disp/uoutp.c b/drivers/gpu/drm/nouveau/nvkm/engine/disp/uoutp.c index 377d0e0cef84..9887b3898505 100644 --- a/drivers/gpu/drm/nouveau/nvkm/engine/disp/uoutp.c +++ b/drivers/gpu/drm/nouveau/nvkm/engine/disp/uoutp.c @@ -253,8 +253,7 @@ nvkm_uoutp_mthd_hdmi(struct nvkm_outp *outp, void *argv, u32 argc) if (!ior->func->hdmi || args->v0.max_ac_packet > 0x1f || - args->v0.rekey > 0x7f || - (args->v0.scdc && !ior->func->hdmi->scdc)) + args->v0.rekey > 0x7f) return -EINVAL; if (!args->v0.enable) { diff --git a/drivers/gpu/drm/nouveau/nvkm/subdev/clk/base.c b/drivers/gpu/drm/nouveau/nvkm/subdev/clk/base.c index 572e63846315..fd2a9d64caba 100644 --- a/drivers/gpu/drm/nouveau/nvkm/subdev/clk/base.c +++ b/drivers/gpu/drm/nouveau/nvkm/subdev/clk/base.c @@ -199,16 +199,18 @@ nvkm_cstate_prog(struct nvkm_clk *clk, struct nvkm_pstate *pstate, int cstatei) } if (volt) { - ret = nvkm_volt_set_id(volt, cstate->voltage, - pstate->base.voltage, clk->temp, -1); - if (ret && ret != -ENODEV) - nvkm_error(subdev, "failed to lower voltage: %d\n", ret); + int err = nvkm_volt_set_id(volt, cstate->voltage, + pstate->base.voltage, clk->temp, -1); + + if (err && err != -ENODEV) + nvkm_error(subdev, "failed to lower voltage: %d\n", err); } if (therm) { - ret = nvkm_therm_cstate(therm, pstate->fanspeed, -1); - if (ret && ret != -ENODEV) - nvkm_error(subdev, "failed to lower fan speed: %d\n", ret); + int err = nvkm_therm_cstate(therm, pstate->fanspeed, -1); + + if (err && err != -ENODEV) + nvkm_error(subdev, "failed to lower fan speed: %d\n", err); } return ret; @@ -473,6 +475,7 @@ static int nvkm_clk_ustate_update(struct nvkm_clk *clk, int req) { struct nvkm_pstate *pstate; + bool found = false; int i = 0; if (!clk->allow_reclock) @@ -480,12 +483,14 @@ nvkm_clk_ustate_update(struct nvkm_clk *clk, int req) if (req != -1 && req != -2) { list_for_each_entry(pstate, &clk->states, head) { - if (pstate->pstate == req) + if (pstate->pstate == req) { + found = true; break; + } i++; } - if (pstate->pstate != req) + if (!found) return -EINVAL; req = i; } diff --git a/drivers/gpu/drm/nouveau/nvkm/subdev/fb/ramnv1a.c b/drivers/gpu/drm/nouveau/nvkm/subdev/fb/ramnv1a.c index 18241c6ba5fa..4d52a158f320 100644 --- a/drivers/gpu/drm/nouveau/nvkm/subdev/fb/ramnv1a.c +++ b/drivers/gpu/drm/nouveau/nvkm/subdev/fb/ramnv1a.c @@ -51,6 +51,8 @@ nv1a_ram_new(struct nvkm_fb *fb, struct nvkm_ram **pram) mib = ((mem >> 4) & 127) + 1; } + pci_dev_put(bridge); + return nvkm_ram_new_(&nv04_ram_func, fb, NVKM_RAM_TYPE_STOLEN, mib * 1024 * 1024, pram); } diff --git a/drivers/gpu/drm/virtio/virtgpu_drv.h b/drivers/gpu/drm/virtio/virtgpu_drv.h index 88fb4be92cf3..5610add1da47 100644 --- a/drivers/gpu/drm/virtio/virtgpu_drv.h +++ b/drivers/gpu/drm/virtio/virtgpu_drv.h @@ -114,6 +114,8 @@ struct virtio_gpu_object { bool dumb; bool created; bool attached; + /* a guest-bound transfer is queued and its mapping not yet synced */ + bool from_host_pending; bool host3d_blob, guest_blob; uint32_t blob_mem, blob_flags; @@ -192,6 +194,9 @@ struct virtio_gpu_vbuffer { struct list_head list; uint32_t seqno; + + /* guest-bound transfer whose shmem backing needs a CPU sync */ + bool sync_for_cpu; }; struct virtio_gpu_output { diff --git a/drivers/gpu/drm/virtio/virtgpu_gem.c b/drivers/gpu/drm/virtio/virtgpu_gem.c index 66c3f6f74e9c..d2f0b8a3f172 100644 --- a/drivers/gpu/drm/virtio/virtgpu_gem.c +++ b/drivers/gpu/drm/virtio/virtgpu_gem.c @@ -45,7 +45,7 @@ static int virtio_gpu_gem_create(struct drm_file *file, ret = drm_gem_handle_create(file, &obj->base.base, &handle); if (ret) { - drm_gem_object_release(&obj->base.base); + drm_gem_object_put(&obj->base.base); return ret; } diff --git a/drivers/gpu/drm/virtio/virtgpu_ioctl.c b/drivers/gpu/drm/virtio/virtgpu_ioctl.c index 3d8e4ccdb7c1..81e70a12b356 100644 --- a/drivers/gpu/drm/virtio/virtgpu_ioctl.c +++ b/drivers/gpu/drm/virtio/virtgpu_ioctl.c @@ -185,7 +185,7 @@ static int virtio_gpu_resource_create_ioctl(struct drm_device *dev, void *data, ret = drm_gem_handle_create(file, obj, &handle); if (ret) { - drm_gem_object_release(obj); + drm_gem_object_put(obj); return ret; } @@ -261,6 +261,27 @@ static int virtio_gpu_transfer_from_host_ioctl(struct drm_device *dev, if (ret != 0) goto err_put_free; + if (virtio_gpu_is_shmem(bo) && virtio_gpu_use_dma_api(vgdev->vdev)) { + /* + * The sync on completion restores the whole mapping, so an + * earlier transfer has to be done before this one snapshots it. + * Otherwise the snapshot predates anything the CPU wrote once + * that transfer's fence signalled, and the later sync would + * discard it. Nothing can add a fence behind our back here, + * since doing so takes the reservation we already hold. + * This writes the pages, so it waits as a writer does. READ + * usage covers existing readers. + */ + long wait = dma_resv_wait_timeout(objs->objs[0]->resv, + DMA_RESV_USAGE_READ, true, + MAX_SCHEDULE_TIMEOUT); + + if (wait < 0) { + ret = wait; + goto err_unlock; + } + } + fence = virtio_gpu_fence_alloc(vgdev, vgdev->fence_drv.context, 0); if (!fence) { ret = -ENOMEM; @@ -320,6 +341,28 @@ static int virtio_gpu_transfer_to_host_ioctl(struct drm_device *dev, void *data, if (ret != 0) goto err_put_free; + /* + * A transfer the other way may have queued without yet syncing + * its mapping. Pushing the guest pages into it now would + * discard what the device wrote there, so wait for that sync: + * it runs before the fence it belongs to is signalled. The + * flag is only set under this reservation, so it cannot appear + * behind our back, and the acquire pairs with the release in + * that sync, so finding it clear means the pages it wrote are + * visible here too. + */ + if (smp_load_acquire(&bo->from_host_pending)) { + long wait = dma_resv_wait_timeout(objs->objs[0]->resv, + DMA_RESV_USAGE_WRITE, + true, + MAX_SCHEDULE_TIMEOUT); + + if (wait < 0) { + ret = wait; + goto err_unlock; + } + } + ret = -ENOMEM; fence = virtio_gpu_fence_alloc(vgdev, vgdev->fence_drv.context, 0); @@ -557,14 +600,14 @@ static int virtio_gpu_resource_create_blob_ioctl(struct drm_device *dev, if (params.blob_flags & VIRTGPU_BLOB_FLAG_USE_CROSS_DEVICE) { ret = virtio_gpu_resource_assign_uuid(vgdev, bo); if (ret) { - drm_gem_object_release(obj); + drm_gem_object_put(obj); return ret; } } ret = drm_gem_handle_create(file, obj, &handle); if (ret) { - drm_gem_object_release(obj); + drm_gem_object_put(obj); return ret; } diff --git a/drivers/gpu/drm/virtio/virtgpu_prime.c b/drivers/gpu/drm/virtio/virtgpu_prime.c index 70b3b836e1c9..c2748378a8ad 100644 --- a/drivers/gpu/drm/virtio/virtgpu_prime.c +++ b/drivers/gpu/drm/virtio/virtgpu_prime.c @@ -310,7 +310,7 @@ struct drm_gem_object *virtgpu_gem_prime_import(struct drm_device *dev, } } - if (!vgdev->has_resource_blob) + if (!vgdev->has_resource_blob || vgdev->has_virgl_3d) return drm_gem_prime_import(dev, buf); bo = kzalloc_obj(*bo); diff --git a/drivers/gpu/drm/virtio/virtgpu_submit.c b/drivers/gpu/drm/virtio/virtgpu_submit.c index 32cb1e4aa425..3d35326dd904 100644 --- a/drivers/gpu/drm/virtio/virtgpu_submit.c +++ b/drivers/gpu/drm/virtio/virtgpu_submit.c @@ -389,10 +389,13 @@ static int virtio_gpu_init_submit(struct virtio_gpu_submit *submit, if ((exbuf->flags & VIRTGPU_EXECBUF_FENCE_FD_OUT) || exbuf->num_out_syncobjs || exbuf->num_bo_handles || - drm_fence_event) + drm_fence_event) { out_fence = virtio_gpu_fence_alloc(vgdev, fence_ctx, ring_idx); - else + if (!out_fence) + return -ENOMEM; + } else { out_fence = NULL; + } if (drm_fence_event) { err = virtio_gpu_fence_event_create(dev, file, out_fence, ring_idx); @@ -538,6 +541,10 @@ int virtio_gpu_execbuffer_ioctl(struct drm_device *dev, void *data, virtio_gpu_process_post_deps(&submit); virtio_gpu_complete_submit(&submit); cleanup: + if (ret && submit.out_fence && submit.out_fence->e) { + drm_event_cancel_free(dev, &submit.out_fence->e->base); + submit.out_fence->e = NULL; + } virtio_gpu_cleanup_submit(&submit); return ret; diff --git a/drivers/gpu/drm/virtio/virtgpu_vq.c b/drivers/gpu/drm/virtio/virtgpu_vq.c index 2b7af8e4e9e6..28682a6d5937 100644 --- a/drivers/gpu/drm/virtio/virtgpu_vq.c +++ b/drivers/gpu/drm/virtio/virtgpu_vq.c @@ -241,6 +241,33 @@ void virtio_gpu_dequeue_ctrl_func(struct work_struct *work) } while (!virtqueue_enable_cb(vgdev->ctrlq.vq)); spin_unlock(&vgdev->ctrlq.qlock); + /* + * Sync guest-bound transfers before signalling anything, so that a + * waiter cannot read the backing pages while what the device wrote is + * still in a bounce buffer. This cannot be folded into the loop below: + * virtio_gpu_fence_event_process() also signals every earlier fence in + * the same context, so any entry there may signal this entry's fence. + */ + list_for_each_entry(entry, &reclaim_list, list) { + if (entry->sync_for_cpu) { + struct virtio_gpu_object *bo = + gem_to_virtio_gpu_obj(entry->objs->objs[0]); + + dma_sync_sgtable_for_cpu(vgdev->vdev->dev.parent, + bo->base.sgt, DMA_FROM_DEVICE); + /* + * Release, so a transfer the other way that skips its + * wait on the strength of this cannot go on to read + * the backing pages before the sync above is visible. + * Nothing orders the two otherwise: where the mapping + * bounces on a coherent device the sync is a plain + * copy, and dma_direct_sync_sg_for_cpu() emits its + * barrier only for the non-coherent case. + */ + smp_store_release(&bo->from_host_pending, false); + } + } + list_for_each_entry(entry, &reclaim_list, list) { resp = (struct virtio_gpu_ctrl_hdr *)entry->resp_buf; @@ -1223,12 +1250,31 @@ void virtio_gpu_cmd_transfer_from_host_3d(struct virtio_gpu_device *vgdev, struct virtio_gpu_object *bo = gem_to_virtio_gpu_obj(objs->objs[0]); struct virtio_gpu_transfer_host_3d *cmd_p; struct virtio_gpu_vbuffer *vbuf; + bool use_dma_api = virtio_gpu_use_dma_api(vgdev->vdev); cmd_p = virtio_gpu_alloc_cmd(vgdev, &vbuf, sizeof(*cmd_p)); memset(cmd_p, 0, sizeof(*cmd_p)); vbuf->objs = objs; + if (virtio_gpu_is_shmem(bo) && use_dma_api) { + /* + * The device writes only the requested box, so prime the + * mapping with the current contents: otherwise the sync on + * completion would hand back whatever a bounce buffer held for + * the regions the device does not touch. + */ + dma_sync_sgtable_for_device(vgdev->vdev->dev.parent, + bo->base.sgt, DMA_TO_DEVICE); + vbuf->sync_for_cpu = true; + /* + * Set under the reservation the caller holds, so a transfer + * the other way cannot miss it and push the guest pages into + * the mapping while the device still owns it. + */ + WRITE_ONCE(bo->from_host_pending, true); + } + cmd_p->hdr.type = cpu_to_le32(VIRTIO_GPU_CMD_TRANSFER_FROM_HOST_3D); cmd_p->hdr.ctx_id = cpu_to_le32(ctx_id); cmd_p->resource_id = cpu_to_le32(bo->hw_res_handle); diff --git a/drivers/gpu/drm/virtio/virtgpu_vram.c b/drivers/gpu/drm/virtio/virtgpu_vram.c index 4ae3cbc35dd3..683ea17454c2 100644 --- a/drivers/gpu/drm/virtio/virtgpu_vram.c +++ b/drivers/gpu/drm/virtio/virtgpu_vram.c @@ -212,16 +212,12 @@ int virtio_gpu_vram_create(struct virtio_gpu_device *vgdev, /* Create fake offset */ ret = drm_gem_create_mmap_offset(obj); - if (ret) { - kfree(vram); - return ret; - } + if (ret) + goto err_release_obj; ret = virtio_gpu_resource_id_get(vgdev, &vram->base.hw_res_handle); - if (ret) { - kfree(vram); - return ret; - } + if (ret) + goto err_release_obj; virtio_gpu_cmd_resource_create_blob(vgdev, &vram->base, params, NULL, 0); @@ -237,6 +233,11 @@ int virtio_gpu_vram_create(struct virtio_gpu_device *vgdev, *bo_ptr = &vram->base; return 0; + +err_release_obj: + drm_gem_object_release(obj); + kfree(vram); + return ret; } void virtio_gpu_vram_map_deferred(struct virtio_gpu_object_vram *vram) diff --git a/drivers/gpu/drm/xe/regs/xe_gt_regs.h b/drivers/gpu/drm/xe/regs/xe_gt_regs.h index 08251c7a1a4b..247a736a54aa 100644 --- a/drivers/gpu/drm/xe/regs/xe_gt_regs.h +++ b/drivers/gpu/drm/xe/regs/xe_gt_regs.h @@ -651,6 +651,7 @@ #define MEM_THERMAL_MASK REG_BIT(2) #define VR_THERMAL_MASK REG_BIT(3) #define ICCMAX_MASK REG_BIT(4) +#define PWRBRK_MASK REG_BIT(5) #define SOC_AVG_THERMAL_MASK REG_BIT(6) #define FASTVMODE_MASK REG_BIT(7) #define PSYS_PL1_MASK REG_BIT(12) diff --git a/drivers/gpu/drm/xe/xe_bo.c b/drivers/gpu/drm/xe/xe_bo.c index 7ed76349075f..2c60e99249f9 100644 --- a/drivers/gpu/drm/xe/xe_bo.c +++ b/drivers/gpu/drm/xe/xe_bo.c @@ -1025,6 +1025,13 @@ static int xe_bo_move(struct ttm_buffer_object *ttm_bo, bool evict, } else { drm_dbg(&xe->drm, "Evict system allocator BO failed=%pe\n", ERR_PTR(ret)); + /* + * The semantic we want upon SVM eviction failure + * because of racing access is keep walking for + * eviction, which is -ENOSPC. + */ + if (ret == -EBUSY) + ret = -ENOSPC; } goto out; diff --git a/drivers/gpu/drm/xe/xe_bo.h b/drivers/gpu/drm/xe/xe_bo.h index 57039cf42ea7..a9e5e1bb8e4f 100644 --- a/drivers/gpu/drm/xe/xe_bo.h +++ b/drivers/gpu/drm/xe/xe_bo.h @@ -9,6 +9,8 @@ #include #include +#include + #include "xe_bo_types.h" #include "xe_ggtt.h" #include "xe_macros.h" @@ -574,6 +576,23 @@ static inline unsigned int xe_sg_segment_size(struct device *dev) struct scatterlist __maybe_unused sg; size_t max = BIT_ULL(sizeof(sg.length) * 8) - 1; + /* + * For Xen PV guests pages aren't contiguous in DMA (machine) address + * space. The DMA API takes care of that both in dma_alloc_* (by + * calling into the hypervisor to make the pages contiguous) and in + * dma_map_* (by bounce buffering). But xe (like i915, see commit + * 78a07fe777c4) ignores the coherency aspects of the DMA API and thus + * can't cope with bounce buffering actually happening, so add a hack + * here to force small allocations and mappings when running in PV + * mode on Xen. + * + * Note this will still break if bounce buffering is required for other + * reasons, like confidential computing hypervisors or PCIe root ports + * with addressing limitations. + */ + if (xen_pv_domain()) + return PAGE_SIZE; + max = min_t(size_t, max, dma_max_mapping_size(dev)); /* diff --git a/drivers/gpu/drm/xe/xe_gt_throttle.c b/drivers/gpu/drm/xe/xe_gt_throttle.c index 1e7e3a31aa69..c0af5484611d 100644 --- a/drivers/gpu/drm/xe/xe_gt_throttle.c +++ b/drivers/gpu/drm/xe/xe_gt_throttle.c @@ -39,7 +39,7 @@ * - ``reason_mem_thermal``: Memory thermal * - ``reason_vr_thermal``: VR thermal * - ``reason_iccmax``: ICCMAX - * - ``reason_ratl``: RATL thermal algorithm + * - ``reason_pwrbrk``: Power brake * - ``reason_soc_avg_thermal``: SoC average temp * - ``reason_fastvmode``: VR is hitting FastVMode * - ``reason_psys_pl1``: PSYS PL1 @@ -200,6 +200,7 @@ static THROTTLE_ATTR_RO(reason_psys_pl1, PSYS_PL1_MASK); static THROTTLE_ATTR_RO(reason_psys_pl2, PSYS_PL2_MASK); static THROTTLE_ATTR_RO(reason_p0_freq, P0_FREQ_MASK); static THROTTLE_ATTR_RO(reason_psys_crit, PSYS_CRIT_MASK); +static THROTTLE_ATTR_RO(reason_pwrbrk, PWRBRK_MASK); static struct attribute *cri_throttle_attrs[] = { /* Common */ @@ -209,12 +210,12 @@ static struct attribute *cri_throttle_attrs[] = { &attr_reason_pl2.attr.attr, &attr_reason_pl4.attr.attr, &attr_reason_prochot.attr.attr, - &attr_reason_ratl.attr.attr, /* CRI */ &attr_reason_vr_thermal.attr.attr, &attr_reason_soc_thermal.attr.attr, &attr_reason_mem_thermal.attr.attr, &attr_reason_iccmax.attr.attr, + &attr_reason_pwrbrk.attr.attr, &attr_reason_soc_avg_thermal.attr.attr, &attr_reason_fastvmode.attr.attr, &attr_reason_psys_pl1.attr.attr, diff --git a/drivers/gpu/drm/xe/xe_guc_ads.c b/drivers/gpu/drm/xe/xe_guc_ads.c index 886bafd31f45..2b42a1bb83b2 100644 --- a/drivers/gpu/drm/xe/xe_guc_ads.c +++ b/drivers/gpu/drm/xe/xe_guc_ads.c @@ -791,7 +791,7 @@ static unsigned int guc_mmio_regset_write(struct xe_guc_ads *ads, } } - if (XE_GT_WA(hwe->gt, 16023105232)) + if (XE_GT_WA(hwe->gt, 16023105232) || XE_GT_WA(hwe->gt, 14025941587)) guc_mmio_regset_write_one(ads, regset_map, RING_IDLEDLY(hwe->mmio_base), count++); diff --git a/drivers/gpu/drm/xe/xe_hw_engine.c b/drivers/gpu/drm/xe/xe_hw_engine.c index 0b193c451a11..ecf691923c7c 100644 --- a/drivers/gpu/drm/xe/xe_hw_engine.c +++ b/drivers/gpu/drm/xe/xe_hw_engine.c @@ -577,28 +577,102 @@ static void hw_engine_init_early(struct xe_gt *gt, struct xe_hw_engine *hwe, xe_reg_whitelist_process_engine(hwe); } +static u32 idledly_floor_ticks(u32 idledly_ns, u32 idledly_units_ps) +{ + return DIV_ROUND_DOWN_ULL((u64)idledly_ns * 1000, idledly_units_ps); +} + static void adjust_idledly(struct xe_hw_engine *hwe) { struct xe_gt *gt = hwe->gt; - u32 idledly, maxcnt; + u32 idledly, idledly_hw, idledly_reg_val, maxcnt; u32 idledly_units_ps = 8 * gt->info.timestamp_base; u32 maxcnt_units_ns = 640; - bool inhibit_switch = 0; + bool inhibit_switch = false; + bool wa_applied = false; + bool clamped_below_maxcnt = false; + + if ((!IS_SRIOV_VF(gt_to_xe(gt)) && XE_GT_WA(gt, 16023105232)) || + XE_GT_WA(gt, 14025941587)) { + u32 mincnt_idledly_ns = 5000; + + /* xe_gt_clock_init() warns and zeroes timestamp_base on unknown crystal clock. */ + if (!idledly_units_ps) + return; - if (!IS_SRIOV_VF(gt_to_xe(hwe->gt)) && XE_GT_WA(gt, 16023105232)) { - idledly = xe_mmio_read32(>->mmio, RING_IDLEDLY(hwe->mmio_base)); + idledly_reg_val = xe_mmio_read32(>->mmio, RING_IDLEDLY(hwe->mmio_base)); maxcnt = xe_mmio_read32(>->mmio, RING_PWRCTX_MAXCNT(hwe->mmio_base)); - inhibit_switch = idledly & INHIBIT_SWITCH_UNTIL_PREEMPTED; - idledly = REG_FIELD_GET(IDLE_DELAY, idledly); - idledly = DIV_ROUND_CLOSEST(idledly * idledly_units_ps, 1000); + inhibit_switch = idledly_reg_val & INHIBIT_SWITCH_UNTIL_PREEMPTED; + idledly = REG_FIELD_GET(IDLE_DELAY, idledly_reg_val); + idledly = DIV_ROUND_CLOSEST_ULL((u64)idledly * idledly_units_ps, 1000); + idledly_hw = idledly; maxcnt = REG_FIELD_GET(IDLE_WAIT_TIME, maxcnt); maxcnt *= maxcnt_units_ns; - if (xe_gt_WARN_ON(gt, idledly >= maxcnt || inhibit_switch)) { - idledly = DIV_ROUND_CLOSEST(((maxcnt - 1) * 1000), - idledly_units_ps); - xe_mmio_write32(>->mmio, RING_IDLEDLY(hwe->mmio_base), idledly); + /* + * Wa_14025941587 is applied before Wa_16023105232, which takes + * priority if the two ever conflict (not expected in practice). + */ + if (XE_GT_WA(gt, 14025941587) && + idledly < mincnt_idledly_ns) { + idledly = mincnt_idledly_ns; + wa_applied = true; + } + + if (XE_GT_WA(gt, 16023105232)) { + /* Clear the inhibit switch without disturbing a valid delay. */ + if (inhibit_switch) { + idledly_reg_val &= ~INHIBIT_SWITCH_UNTIL_PREEMPTED; + wa_applied = true; + } + + /* Warn only on the value read from hardware. */ + xe_gt_WARN_ON(gt, idledly_hw >= maxcnt); + + if (idledly >= maxcnt) { + /* maxcnt may be 0 if IDLE_WAIT_TIME is unprogrammed. */ + idledly = maxcnt ? maxcnt - 1 : 0; + clamped_below_maxcnt = true; + wa_applied = true; + } + } + + if (wa_applied) { + u32 idledly_ticks; + + /* + * Wa_16023105232 requires idledly < maxcnt, so floor + * that clamp; otherwise round up to guarantee the + * Wa_14025941587 minimum survives tick quantization. + */ + if (clamped_below_maxcnt) + idledly_ticks = idledly_floor_ticks(idledly, idledly_units_ps); + else + idledly_ticks = DIV_ROUND_UP_ULL((u64)idledly * 1000, + idledly_units_ps); + + /* + * Tick quantization can still push the rounded-up value + * to/above maxcnt; re-floor here so Wa_16023105232 keeps + * priority even in that case. + */ + if (!clamped_below_maxcnt && XE_GT_WA(gt, 16023105232) && + (u64)idledly_ticks * idledly_units_ps >= (u64)maxcnt * 1000) { + xe_gt_dbg(gt, "idledly %s: %u ticks would exceed maxcnt=%u, so flooring\n", + hwe->name, idledly_ticks, maxcnt); + idledly = maxcnt ? maxcnt - 1 : 0; + idledly_ticks = idledly_floor_ticks(idledly, idledly_units_ps); + } + + idledly_reg_val &= ~IDLE_DELAY; + idledly_reg_val |= REG_FIELD_PREP(IDLE_DELAY, idledly_ticks); + xe_gt_dbg(gt, "idledly %s: set %u max=%u inh=%u ts=%u\n", + hwe->name, idledly, maxcnt, + !!inhibit_switch, gt->info.timestamp_base); + xe_mmio_write32(>->mmio, + RING_IDLEDLY(hwe->mmio_base), + idledly_reg_val); } } } diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c index 67819deb45e3..84ced696f46b 100644 --- a/drivers/gpu/drm/xe/xe_vm.c +++ b/drivers/gpu/drm/xe/xe_vm.c @@ -1933,21 +1933,13 @@ void xe_vm_close_and_put(struct xe_vm *vm) vma->gpuva.flags |= XE_VMA_DESTROYED; } - /* - * All vm operations will add shared fences to resv. - * The only exception is eviction for a shared object, - * but even so, the unbind when evicted would still - * install a fence to resv. Hence it's safe to - * destroy the pagetables immediately. - */ - xe_vm_free_scratch(vm); - xe_vm_pt_destroy(vm); xe_vm_unlock(vm); /* - * VM is now dead, cannot re-add nodes to vm->vmas if it's NULL - * Since we hold a refcount to the bo, we can remove and free - * the members safely without locking. + * Unlink and destroy all contested external-BO VMAs before destroying + * the page tables. Otherwise, concurrent eviction holding only bo->resv + * can walk the BO's VMAs and attempt to invalidate/zap page tables that + * have already been freed. */ list_for_each_entry_safe(vma, next_vma, &contested, combined_links.destroy) { @@ -1955,6 +1947,11 @@ void xe_vm_close_and_put(struct xe_vm *vm) xe_vma_destroy_unlocked(vma); } + xe_vm_lock(vm, false); + xe_vm_free_scratch(vm); + xe_vm_pt_destroy(vm); + xe_vm_unlock(vm); + xe_svm_fini(vm); up_write(&vm->lock); diff --git a/drivers/gpu/drm/xe/xe_wa_oob.rules b/drivers/gpu/drm/xe/xe_wa_oob.rules index 132a41859ab7..00c58b3ca9c2 100644 --- a/drivers/gpu/drm/xe/xe_wa_oob.rules +++ b/drivers/gpu/drm/xe/xe_wa_oob.rules @@ -65,3 +65,5 @@ 14025883347 MEDIA_VERSION_RANGE(1301, 3500) GRAPHICS_VERSION_RANGE(2004, 3005) +14025941587 GRAPHICS_VERSION_RANGE(2001, 3511), FUNC(xe_rtp_match_not_sriov_vf) + MEDIA_VERSION_RANGE(1301, 3503), FUNC(xe_rtp_match_not_sriov_vf) diff --git a/drivers/hid/amd-sfh-hid/amd_sfh_common.h b/drivers/hid/amd-sfh-hid/amd_sfh_common.h index 78f830c133e5..bd8dc16feb61 100644 --- a/drivers/hid/amd-sfh-hid/amd_sfh_common.h +++ b/drivers/hid/amd-sfh-hid/amd_sfh_common.h @@ -12,11 +12,15 @@ #include #include +#include #include "amd_sfh_hid.h" #define PCI_DEVICE_ID_AMD_MP2 0x15E4 #define PCI_DEVICE_ID_AMD_MP2_1_1 0x164A +/* The BAR 2 size must cover the highest register offset (0x10958) */ +#define AMD_SFH_MIN_BAR_SIZE SZ_128K + #define AMD_C2P_MSG(regno) (0x10500 + ((regno) * 4)) #define AMD_P2C_MSG(regno) (0x10680 + ((regno) * 4)) diff --git a/drivers/hid/amd-sfh-hid/amd_sfh_pcie.c b/drivers/hid/amd-sfh-hid/amd_sfh_pcie.c index 4b81cebdc335..039b6ac327d3 100644 --- a/drivers/hid/amd-sfh-hid/amd_sfh_pcie.c +++ b/drivers/hid/amd-sfh-hid/amd_sfh_pcie.c @@ -451,6 +451,16 @@ static int amd_mp2_pci_probe(struct pci_dev *pdev, const struct pci_device_id *i if (rc) return rc; + if (!(pci_resource_flags(pdev, 2) & IORESOURCE_MEM)) { + dev_err(&pdev->dev, "BAR 2 is not IORESOURCE_MEM\n"); + return -ENODEV; + } + + if (pci_resource_len(pdev, 2) < AMD_SFH_MIN_BAR_SIZE) { + dev_err(&pdev->dev, "BAR 2 is too small\n"); + return -EINVAL; + } + rc = pcim_iomap_regions(pdev, BIT(2), DRIVER_NAME); if (rc) return rc; diff --git a/drivers/hid/bpf/hid_bpf_dispatch.c b/drivers/hid/bpf/hid_bpf_dispatch.c index 536f6d01fd14..b1de1dd0f21d 100644 --- a/drivers/hid/bpf/hid_bpf_dispatch.c +++ b/drivers/hid/bpf/hid_bpf_dispatch.c @@ -359,7 +359,7 @@ hid_bpf_release_context(struct hid_bpf_ctx *ctx) static int __hid_bpf_hw_check_params(struct hid_bpf_ctx *ctx, __u8 *buf, size_t *buf__sz, - enum hid_report_type rtype) + enum hid_report_type rtype, bool hw_request) { struct hid_report_enum *report_enum; struct hid_report *report; @@ -388,6 +388,10 @@ __hid_bpf_hw_check_params(struct hid_bpf_ctx *ctx, __u8 *buf, size_t *buf__sz, report_len = hid_report_len(report); + /* unnumbered reports need to have a report ID reserved in the first byte */ + if (hw_request && report_enum->numbered == 0) + report_len += 1; + if (*buf__sz > report_len) *buf__sz = report_len; @@ -420,7 +424,7 @@ hid_bpf_hw_request(struct hid_bpf_ctx *ctx, __u8 *buf, size_t buf__sz, return -EDEADLOCK; /* check arguments */ - ret = __hid_bpf_hw_check_params(ctx, buf, &size, rtype); + ret = __hid_bpf_hw_check_params(ctx, buf, &size, rtype, true); if (ret) return ret; @@ -480,7 +484,7 @@ hid_bpf_hw_output_report(struct hid_bpf_ctx *ctx, __u8 *buf, size_t buf__sz) return -EDEADLOCK; /* check arguments */ - ret = __hid_bpf_hw_check_params(ctx, buf, &size, HID_OUTPUT_REPORT); + ret = __hid_bpf_hw_check_params(ctx, buf, &size, HID_OUTPUT_REPORT, true); if (ret) return ret; @@ -506,7 +510,7 @@ __hid_bpf_input_report(struct hid_bpf_ctx *ctx, enum hid_report_type type, u8 *b return -EDEADLOCK; /* check arguments */ - ret = __hid_bpf_hw_check_params(ctx, buf, &size, type); + ret = __hid_bpf_hw_check_params(ctx, buf, &size, type, false); if (ret) return ret; diff --git a/drivers/hid/hid-alps.c b/drivers/hid/hid-alps.c index 67179e3fe39b..0556cb5645eb 100644 --- a/drivers/hid/hid-alps.c +++ b/drivers/hid/hid-alps.c @@ -407,6 +407,8 @@ static int u1_raw_event(struct alps_dev *hdata, u8 *data, int size) return 1; case U1_SP_ABSOLUTE_REPORT_ID: + if (!hdata->input2) + return 0; sp_x = get_unaligned_le16(data+2); sp_y = get_unaligned_le16(data+4); @@ -738,7 +740,6 @@ static int alps_input_configured(struct hid_device *hdev, struct hid_input *hi) goto exit; } - data->input2 = input2; input2->phys = input->phys; input2->name = "DualPoint Stick"; input2->id.bustype = BUS_I2C; @@ -762,11 +763,12 @@ static int alps_input_configured(struct hid_device *hdev, struct hid_input *hi) __set_bit(INPUT_PROP_POINTER, input2->propbit); __set_bit(INPUT_PROP_POINTING_STICK, input2->propbit); - if (input_register_device(data->input2)) { + if (input_register_device(input2)) { input_free_device(input2); ret = -ENOENT; goto exit; } + data->input2 = input2; } exit: @@ -823,6 +825,24 @@ static int alps_probe(struct hid_device *hdev, const struct hid_device_id *id) return 0; } +static void alps_remove(struct hid_device *hdev) +{ + struct alps_dev *data = hid_get_drvdata(hdev); + + /* + * input2 ("DualPoint Stick") is allocated separately and is not + * tracked in hdev->inputs, so the default remove path + * (hid_hw_stop -> hidinput_disconnect) does not unregister it. + * + * Stop the device first so that no URB callback can touch input2 + * while it is being unregistered, then drop it explicitly. + */ + hid_hw_stop(hdev); + + if (data->input2) + input_unregister_device(data->input2); +} + static const struct hid_device_id alps_id[] = { { HID_DEVICE(HID_BUS_ANY, HID_GROUP_ANY, USB_VENDOR_ID_ALPS_JP, HID_DEVICE_ID_ALPS_U1_DUAL) }, @@ -845,6 +865,7 @@ static struct hid_driver alps_driver = { .input_configured = alps_input_configured, .resume = pm_ptr(alps_post_resume), .reset_resume = pm_ptr(alps_post_reset), + .remove = alps_remove, }; module_hid_driver(alps_driver); diff --git a/drivers/hid/hid-ids.h b/drivers/hid/hid-ids.h index 1059922baaac..dd31757601b6 100644 --- a/drivers/hid/hid-ids.h +++ b/drivers/hid/hid-ids.h @@ -1283,6 +1283,9 @@ #define USB_DEVICE_ID_SAMSUNG_WIRELESS_UNIVERSAL_KBD 0xa006 #define USB_DEVICE_ID_SAMSUNG_WIRELESS_MULTI_HOGP_KBD 0xa064 +#define USB_VENDOR_ID_SDINNOVATION 0x36ae +#define USB_DEVICE_ID_SDINNOVATION_GAMING_KBD 0xfeab + #define USB_VENDOR_ID_SEMICO 0x1a2c #define USB_DEVICE_ID_SEMICO_USB_KEYKOARD 0x0023 #define USB_DEVICE_ID_SEMICO_USB_KEYKOARD2 0x0027 diff --git a/drivers/hid/hid-multitouch.c b/drivers/hid/hid-multitouch.c index 571166a769b9..e9da59e4bf79 100644 --- a/drivers/hid/hid-multitouch.c +++ b/drivers/hid/hid-multitouch.c @@ -79,6 +79,7 @@ MODULE_LICENSE("GPL"); #define MT_QUIRK_APPLE_TOUCHBAR BIT(23) #define MT_QUIRK_YOGABOOK9I BIT(24) #define MT_QUIRK_KEEP_LATENCY_ON_CLOSE BIT(25) +#define MT_QUIRK_IGNORE_FEATURE_ID_MISMATCH BIT(26) #define MT_INPUTMODE_TOUCHSCREEN 0x02 #define MT_INPUTMODE_TOUCHPAD 0x03 @@ -235,6 +236,7 @@ static void mt_post_parse(struct mt_device *td, struct mt_application *app); #define MT_CLS_APPLE_TOUCHBAR 0x0114 #define MT_CLS_YOGABOOK9I 0x0115 #define MT_CLS_EGALAX_P80H84 0x0116 +#define MT_CLS_ASUS_ROG_Z13_FOLIO 0x0117 #define MT_CLS_SIS 0x0457 #define MT_DEFAULT_MAXCONTACT 10 @@ -405,6 +407,16 @@ static const struct mt_class mt_classes[] = { .quirks = MT_QUIRK_ALWAYS_VALID | MT_QUIRK_CONTACT_CNT_ACCURATE | MT_QUIRK_ASUS_CUSTOM_UP }, + { .name = MT_CLS_ASUS_ROG_Z13_FOLIO, + .quirks = MT_QUIRK_ALWAYS_VALID | + MT_QUIRK_IGNORE_DUPLICATES | + MT_QUIRK_HOVERING | + MT_QUIRK_CONTACT_CNT_ACCURATE | + MT_QUIRK_STICKY_FINGERS | + MT_QUIRK_WIN8_PTP_BUTTONS | + MT_QUIRK_CONFIDENCE | + MT_QUIRK_IGNORE_FEATURE_ID_MISMATCH, + .export_all_inputs = true }, { .name = MT_CLS_VTL, .quirks = MT_QUIRK_ALWAYS_VALID | MT_QUIRK_CONTACT_CNT_ACCURATE | @@ -507,6 +519,7 @@ static const struct attribute_group mt_attribute_group = { static void mt_get_feature(struct hid_device *hdev, struct hid_report *report) { + struct mt_device *td = hid_get_drvdata(hdev); int ret; u32 size = hid_report_len(report); u8 *buf; @@ -528,8 +541,14 @@ static void mt_get_feature(struct hid_device *hdev, struct hid_report *report) dev_warn(&hdev->dev, "failed to fetch feature %d\n", report->id); } else { - /* The report ID in the request and the response should match */ - if (report->id != buf[0]) { + /* + * The report ID in the request and the response should match. + * Some firmware (e.g. the ASUS ROG Z13 Folio + * touchpad) returns a mismatched ID on this specific fetch; + * tolerate it only for devices explicitly flagged as such. + */ + if (report->id != buf[0] && + !(td->mtclass.quirks & MT_QUIRK_IGNORE_FEATURE_ID_MISMATCH)) { hid_err(hdev, "Returned feature report did not match the request\n"); goto free; } @@ -2726,6 +2745,12 @@ static const struct hid_device_id mt_devices[] = { HID_DEVICE(BUS_I2C, HID_GROUP_MULTITOUCH_WIN_8, I2C_VENDOR_ID_HANTICK, I2C_PRODUCT_ID_HANTICK_5288) }, + /* Asus ROG Flow Z13 (2025) GZ302EA keyboard-cover touchpad */ + { .driver_data = MT_CLS_ASUS_ROG_Z13_FOLIO, + HID_DEVICE(BUS_USB, HID_GROUP_MULTITOUCH_WIN_8, + USB_VENDOR_ID_ASUSTEK, + USB_DEVICE_ID_ASUSTEK_ROG_Z13_FOLIO) }, + /* Generic MT device */ { HID_DEVICE(HID_BUS_ANY, HID_GROUP_MULTITOUCH, HID_ANY_ID, HID_ANY_ID) }, diff --git a/drivers/hid/hid-oxp.c b/drivers/hid/hid-oxp.c index 20a54f337220..d8fb6a69d40d 100644 --- a/drivers/hid/hid-oxp.c +++ b/drivers/hid/hid-oxp.c @@ -1552,9 +1552,9 @@ static int oxp_hid_probe(struct hid_device *hdev, static void oxp_hid_remove(struct hid_device *hdev) { - cancel_delayed_work(&drvdata.oxp_rgb_queue); - cancel_delayed_work(&drvdata.oxp_btn_queue); - cancel_delayed_work(&drvdata.oxp_mcu_init); + cancel_delayed_work_sync(&drvdata.oxp_rgb_queue); + cancel_delayed_work_sync(&drvdata.oxp_btn_queue); + cancel_delayed_work_sync(&drvdata.oxp_mcu_init); hid_hw_close(hdev); hid_hw_stop(hdev); } diff --git a/drivers/hid/hid-quirks.c b/drivers/hid/hid-quirks.c index 57d8efdd9b89..dcfa8278447f 100644 --- a/drivers/hid/hid-quirks.c +++ b/drivers/hid/hid-quirks.c @@ -183,6 +183,7 @@ static const struct hid_device_id hid_quirks[] = { { HID_USB_DEVICE(USB_VENDOR_ID_SAITEK, USB_DEVICE_ID_SAITEK_X52_2), HID_QUIRK_INCREMENT_USAGE_ON_DUPLICATE }, { HID_USB_DEVICE(USB_VENDOR_ID_SAITEK, USB_DEVICE_ID_SAITEK_X52_PRO), HID_QUIRK_INCREMENT_USAGE_ON_DUPLICATE }, { HID_USB_DEVICE(USB_VENDOR_ID_SAITEK, USB_DEVICE_ID_SAITEK_X65), HID_QUIRK_INCREMENT_USAGE_ON_DUPLICATE }, + { HID_USB_DEVICE(USB_VENDOR_ID_SDINNOVATION, USB_DEVICE_ID_SDINNOVATION_GAMING_KBD), HID_QUIRK_ALWAYS_POLL }, { HID_USB_DEVICE(USB_VENDOR_ID_SEMICO, USB_DEVICE_ID_SEMICO_USB_KEYKOARD2), HID_QUIRK_NO_INIT_REPORTS }, { HID_USB_DEVICE(USB_VENDOR_ID_SEMICO, USB_DEVICE_ID_SEMICO_USB_KEYKOARD), HID_QUIRK_NO_INIT_REPORTS }, { HID_USB_DEVICE(USB_VENDOR_ID_SENNHEISER, USB_DEVICE_ID_SENNHEISER_BTD500USB), HID_QUIRK_NOGET }, @@ -422,7 +423,7 @@ static const struct hid_device_id hid_have_special_driver[] = { #endif #if IS_ENABLED(CONFIG_HID_ELECOM) { HID_BLUETOOTH_DEVICE(USB_VENDOR_ID_ELECOM, USB_DEVICE_ID_ELECOM_BM084) }, - { HID_BLUETOOTH_DEVICE(USB_VENDOR_ID_ELECOM, USB_DEVICE_ID_ELECOM_M_XGL20DLBK) }, + { HID_USB_DEVICE(USB_VENDOR_ID_ELECOM, USB_DEVICE_ID_ELECOM_M_XGL20DLBK) }, { HID_BLUETOOTH_DEVICE(USB_VENDOR_ID_ELECOM, USB_DEVICE_ID_ELECOM_M_HT1MRBK_01AC) }, { HID_USB_DEVICE(USB_VENDOR_ID_ELECOM, USB_DEVICE_ID_ELECOM_M_XT3URBK_00FB) }, { HID_USB_DEVICE(USB_VENDOR_ID_ELECOM, USB_DEVICE_ID_ELECOM_M_XT3URBK_018F) }, diff --git a/drivers/hid/hid-winwing.c b/drivers/hid/hid-winwing.c index 9cd25a77999e..19b92c2c6579 100644 --- a/drivers/hid/hid-winwing.c +++ b/drivers/hid/hid-winwing.c @@ -315,7 +315,8 @@ static void winwing_haptic_rumble_cb(struct work_struct *work) static int winwing_play_effect(struct input_dev *dev, void *context, struct ff_effect *effect) { - struct winwing_drv_data *data = (struct winwing_drv_data *) context; + struct hid_device *hdev = input_get_drvdata(dev); + struct winwing_drv_data *data = hid_get_drvdata(hdev); if (effect->type != FF_RUMBLE) return 0; @@ -342,7 +343,12 @@ static int winwing_init_ff(struct hid_device *hdev, struct hid_input *hidinput) input_set_capability(hidinput->input, EV_FF, FF_RUMBLE); - return input_ff_create_memless(hidinput->input, data, + /* + * input_ff_create_memless() takes ownership of the context pointer + * and frees it on teardown; do not hand it the devm-managed drvdata. + * winwing_play_effect() fetches it from the input device instead. + */ + return input_ff_create_memless(hidinput->input, NULL, winwing_play_effect); } diff --git a/drivers/hid/wacom_sys.c b/drivers/hid/wacom_sys.c index 0eafa483b7f7..40770affdbde 100644 --- a/drivers/hid/wacom_sys.c +++ b/drivers/hid/wacom_sys.c @@ -113,8 +113,9 @@ static int wacom_wac_pen_serial_enforce(struct hid_device *hdev, /* Queue events which have invalid tool type or serial number */ for (i = 0; i < report->maxfield; i++) { - for (j = 0; j < report->field[i]->maxusage; j++) { - struct hid_field *field = report->field[i]; + struct hid_field *field = report->field[i]; + + for (j = 0; j < field->report_count; j++) { struct hid_usage *usage = &field->usage[j]; unsigned int equivalent_usage = wacom_equivalent_usage(usage->hid); unsigned int offset; diff --git a/drivers/i2c/busses/i2c-qcom-geni.c b/drivers/i2c/busses/i2c-qcom-geni.c index 9a02e57c5a4b..de5bc4b8a3a5 100644 --- a/drivers/i2c/busses/i2c-qcom-geni.c +++ b/drivers/i2c/busses/i2c-qcom-geni.c @@ -77,6 +77,14 @@ enum geni_i2c_err_code { #define XFER_TIMEOUT HZ #define RST_TIMEOUT HZ +#define GENI_SE_CLK_32MHZ (32 * HZ_PER_MHZ) +#define GENI_SE_CLK_19P2MHZ 19200000UL + +struct geni_i2c_desc { + bool no_dma_support; + unsigned int tx_fifo_depth; +}; + #define QCOM_I2C_MIN_NUM_OF_MSGS_MULTI_DESC 2 /** @@ -107,9 +115,9 @@ struct geni_i2c_dev { int cur_wr; int cur_rd; spinlock_t lock; - struct clk *core_clk; u32 clk_freq_out; const struct geni_i2c_clk_fld *clk_fld; + u32 clk_idx; void *dma_buf; size_t xfer_len; dma_addr_t dma_addr; @@ -121,13 +129,7 @@ struct geni_i2c_dev { bool is_tx_multi_desc_xfer; u32 num_msgs; struct geni_i2c_gpi_multi_desc_xfer i2c_multi_desc_config; -}; - -struct geni_i2c_desc { - bool has_core_clk; - char *icc_ddr; - bool no_dma_support; - unsigned int tx_fifo_depth; + const struct geni_i2c_desc *dev_data; }; struct geni_i2c_err_log { @@ -186,19 +188,44 @@ static const struct geni_i2c_clk_fld geni_i2c_clk_map_32mhz[] = { static int geni_i2c_clk_map_idx(struct geni_i2c_dev *gi2c) { const struct geni_i2c_clk_fld *itr; + unsigned long res_freq; - if (clk_get_rate(gi2c->se.clk) == 32 * HZ_PER_MHZ) + /* + * Frequency counter tables are calibrated for a specific source + * clock frequency and are not valid for any multiple of it + * (e.g. 64 MHz, 128 MHz). + * Use exact=true and verify res_freq matches req_freq literally + * to reject harmonics: a 64 MHz clock that divides evenly to + * 32 MHz would pass exact matching but produce double the intended + * I2C frequency with these counter values. + */ + if (!geni_se_clk_freq_match(&gi2c->se, GENI_SE_CLK_32MHZ, + &gi2c->clk_idx, &res_freq, true) && + res_freq == GENI_SE_CLK_32MHZ) { itr = geni_i2c_clk_map_32mhz; - else + } else if (!geni_se_clk_freq_match(&gi2c->se, GENI_SE_CLK_19P2MHZ, + &gi2c->clk_idx, &res_freq, true) && + res_freq == GENI_SE_CLK_19P2MHZ) { itr = geni_i2c_clk_map_19p2mhz; + } else { + dev_err(gi2c->se.dev, + "Unsupported SE source clock: must be exactly 32 MHz or 19.2 MHz\n"); + return -EINVAL; + } while (itr->clk_freq_out != 0) { if (itr->clk_freq_out == gi2c->clk_freq_out) { gi2c->clk_fld = itr; + dev_dbg(gi2c->se.dev, + "I2C clk selected: freq: %u Hz, clk_idx: %u\n", + gi2c->clk_freq_out, gi2c->clk_idx); return 0; } itr++; } + + dev_err(gi2c->se.dev, "Unsupported I2C output frequency %u Hz\n", gi2c->clk_freq_out); + return -EINVAL; } @@ -207,7 +234,7 @@ static void qcom_geni_i2c_conf(struct geni_i2c_dev *gi2c) const struct geni_i2c_clk_fld *itr = gi2c->clk_fld; u32 val; - writel_relaxed(0, gi2c->se.base + SE_GENI_CLK_SEL); + writel_relaxed(gi2c->clk_idx, gi2c->se.base + SE_GENI_CLK_SEL); val = (itr->clk_div << CLK_DIV_SHFT) | SER_CLK_EN; writel_relaxed(val, gi2c->se.base + GENI_SER_M_CLK_CFG); @@ -944,15 +971,6 @@ static const struct i2c_algorithm geni_i2c_algo = { .functionality = geni_i2c_func, }; -#ifdef CONFIG_ACPI -static const struct acpi_device_id geni_i2c_acpi_match[] = { - { "QCOM0220"}, - { "QCOM0411" }, - { } -}; -MODULE_DEVICE_TABLE(acpi, geni_i2c_acpi_match); -#endif - static void release_gpi_dma(struct geni_i2c_dev *gi2c) { if (gi2c->rx_c) @@ -990,13 +1008,93 @@ static int setup_gpi_dma(struct geni_i2c_dev *gi2c) return ret; } +static int geni_i2c_init(struct geni_i2c_dev *gi2c) +{ + u32 proto, tx_depth; + bool fifo_disable; + int ret; + + ret = pm_runtime_resume_and_get(gi2c->se.dev); + if (ret < 0) { + dev_err(gi2c->se.dev, "error turning on device :%d\n", ret); + return ret; + } + + proto = geni_se_read_proto(&gi2c->se); + if (proto == GENI_SE_INVALID_PROTO) { + ret = geni_load_se_firmware(&gi2c->se, GENI_SE_I2C); + if (ret) { + dev_err_probe(gi2c->se.dev, ret, "i2c firmware load failed ret: %d\n", ret); + goto err; + } + } else if (proto != GENI_SE_I2C) { + ret = dev_err_probe(gi2c->se.dev, -ENXIO, "Invalid proto %d\n", proto); + goto err; + } + + if (gi2c->dev_data->no_dma_support) { + fifo_disable = false; + gi2c->no_dma = true; + } else { + fifo_disable = readl_relaxed(gi2c->se.base + GENI_IF_DISABLE_RO) & FIFO_IF_DISABLE; + } + + if (fifo_disable) { + /* FIFO is disabled, so we can only use GPI DMA */ + gi2c->gpi_mode = true; + ret = setup_gpi_dma(gi2c); + if (ret) + goto err; + + dev_dbg(gi2c->se.dev, "Using GPI DMA mode for I2C\n"); + } else { + gi2c->gpi_mode = false; + tx_depth = geni_se_get_tx_fifo_depth(&gi2c->se); + + /* I2C Master Hub Serial Elements doesn't have the HW_PARAM_0 register */ + if (!tx_depth && gi2c->se.core_clk) + tx_depth = gi2c->dev_data->tx_fifo_depth; + + if (!tx_depth) { + ret = dev_err_probe(gi2c->se.dev, -EINVAL, + "Invalid TX FIFO depth\n"); + goto err; + } + + gi2c->tx_wm = tx_depth - 1; + geni_se_init(&gi2c->se, gi2c->tx_wm, tx_depth); + geni_se_config_packing(&gi2c->se, BITS_PER_BYTE, + PACKING_BYTES_PW, true, true, true); + + dev_dbg(gi2c->se.dev, "i2c fifo/se-dma mode. fifo depth:%d\n", tx_depth); + } + +err: + pm_runtime_put(gi2c->se.dev); + return ret; +} + +static int geni_i2c_resources_init(struct geni_i2c_dev *gi2c) +{ + int ret; + + ret = geni_se_resources_init(&gi2c->se); + if (ret) + return ret; + + ret = geni_i2c_clk_map_idx(gi2c); + if (ret) + return ret; + + return geni_icc_set_bw_ab(&gi2c->se, GENI_DEFAULT_BW, GENI_DEFAULT_BW, + Bps_to_icc(gi2c->clk_freq_out)); +} + static int geni_i2c_probe(struct platform_device *pdev) { struct geni_i2c_dev *gi2c; - u32 proto, tx_depth, fifo_disable; int ret; struct device *dev = &pdev->dev; - const struct geni_i2c_desc *desc = NULL; gi2c = devm_kzalloc(dev, sizeof(*gi2c), GFP_KERNEL); if (!gi2c) @@ -1008,17 +1106,9 @@ static int geni_i2c_probe(struct platform_device *pdev) if (IS_ERR(gi2c->se.base)) return PTR_ERR(gi2c->se.base); - desc = device_get_match_data(&pdev->dev); - - if (desc && desc->has_core_clk) { - gi2c->core_clk = devm_clk_get(dev, "core"); - if (IS_ERR(gi2c->core_clk)) - return PTR_ERR(gi2c->core_clk); - } - - gi2c->se.clk = devm_clk_get(dev, "se"); - if (IS_ERR(gi2c->se.clk) && !has_acpi_companion(dev)) - return PTR_ERR(gi2c->se.clk); + gi2c->dev_data = device_get_match_data(&pdev->dev); + if (!gi2c->dev_data) + return -EINVAL; ret = device_property_read_u32(dev, "clock-frequency", &gi2c->clk_freq_out); @@ -1034,16 +1124,15 @@ static int geni_i2c_probe(struct platform_device *pdev) if (gi2c->irq < 0) return gi2c->irq; - ret = geni_i2c_clk_map_idx(gi2c); - if (ret) - return dev_err_probe(dev, ret, "Invalid clk frequency %d Hz\n", - gi2c->clk_freq_out); - gi2c->adap.algo = &geni_i2c_algo; init_completion(&gi2c->done); spin_lock_init(&gi2c->lock); platform_set_drvdata(pdev, gi2c); + ret = geni_i2c_resources_init(gi2c); + if (ret) + return ret; + /* Keep interrupts disabled initially to allow for low-power modes */ ret = devm_request_irq(dev, gi2c->irq, geni_i2c_irq, IRQF_NO_AUTOEN, dev_name(dev), gi2c); @@ -1056,118 +1145,27 @@ static int geni_i2c_probe(struct platform_device *pdev) gi2c->adap.dev.of_node = dev->of_node; strscpy(gi2c->adap.name, "Geni-I2C", sizeof(gi2c->adap.name)); - ret = geni_icc_get(&gi2c->se, desc ? desc->icc_ddr : "qup-memory"); - if (ret) - return ret; - /* - * Set the bus quota for core and cpu to a reasonable value for - * register access. - * Set quota for DDR based on bus speed. - */ - gi2c->se.icc_paths[GENI_TO_CORE].avg_bw = GENI_DEFAULT_BW; - gi2c->se.icc_paths[CPU_TO_GENI].avg_bw = GENI_DEFAULT_BW; - if (!desc || desc->icc_ddr) - gi2c->se.icc_paths[GENI_TO_DDR].avg_bw = Bps_to_icc(gi2c->clk_freq_out); - - ret = geni_icc_set_bw(&gi2c->se); - if (ret) - return ret; - - ret = clk_prepare_enable(gi2c->core_clk); - if (ret) - return ret; - - ret = geni_se_resources_on(&gi2c->se); - if (ret) { - dev_err_probe(dev, ret, "Error turning on resources\n"); - goto err_clk; - } - proto = geni_se_read_proto(&gi2c->se); - if (proto == GENI_SE_INVALID_PROTO) { - ret = geni_load_se_firmware(&gi2c->se, GENI_SE_I2C); - if (ret) { - dev_err_probe(dev, ret, "i2c firmware load failed ret: %d\n", ret); - goto err_resources; - } - } else if (proto != GENI_SE_I2C) { - ret = dev_err_probe(dev, -ENXIO, "Invalid proto %d\n", proto); - goto err_resources; - } - - if (desc && desc->no_dma_support) { - fifo_disable = false; - gi2c->no_dma = true; - } else { - fifo_disable = readl_relaxed(gi2c->se.base + GENI_IF_DISABLE_RO) & FIFO_IF_DISABLE; - } - - if (fifo_disable) { - /* FIFO is disabled, so we can only use GPI DMA */ - gi2c->gpi_mode = true; - ret = setup_gpi_dma(gi2c); - if (ret) - goto err_resources; - - dev_dbg(dev, "Using GPI DMA mode for I2C\n"); - } else { - gi2c->gpi_mode = false; - tx_depth = geni_se_get_tx_fifo_depth(&gi2c->se); - - /* I2C Master Hub Serial Elements doesn't have the HW_PARAM_0 register */ - if (!tx_depth && desc) - tx_depth = desc->tx_fifo_depth; - - if (!tx_depth) { - ret = dev_err_probe(dev, -EINVAL, - "Invalid TX FIFO depth\n"); - goto err_resources; - } - - gi2c->tx_wm = tx_depth - 1; - geni_se_init(&gi2c->se, gi2c->tx_wm, tx_depth); - geni_se_config_packing(&gi2c->se, BITS_PER_BYTE, - PACKING_BYTES_PW, true, true, true); - - dev_dbg(dev, "i2c fifo/se-dma mode. fifo depth:%d\n", tx_depth); - } - - clk_disable_unprepare(gi2c->core_clk); - ret = geni_se_resources_off(&gi2c->se); - if (ret) { - dev_err_probe(dev, ret, "Error turning off resources\n"); - goto err_dma; - } - - ret = geni_icc_disable(&gi2c->se); - if (ret) - goto err_dma; - pm_runtime_set_suspended(gi2c->se.dev); pm_runtime_set_autosuspend_delay(gi2c->se.dev, I2C_AUTO_SUSPEND_DELAY); pm_runtime_use_autosuspend(gi2c->se.dev); pm_runtime_enable(gi2c->se.dev); + ret = geni_i2c_init(gi2c); + if (ret < 0) { + pm_runtime_disable(gi2c->se.dev); + return ret; + } + ret = i2c_add_adapter(&gi2c->adap); if (ret) { + release_gpi_dma(gi2c); dev_err_probe(dev, ret, "Error adding i2c adapter\n"); pm_runtime_disable(gi2c->se.dev); - goto err_dma; + return ret; } dev_dbg(dev, "Geni-I2C adaptor successfully added\n"); - return ret; - -err_resources: - geni_se_resources_off(&gi2c->se); -err_clk: - clk_disable_unprepare(gi2c->core_clk); - - return ret; - -err_dma: - release_gpi_dma(gi2c); - return ret; } @@ -1200,7 +1198,7 @@ static int __maybe_unused geni_i2c_runtime_suspend(struct device *dev) return ret; } - clk_disable_unprepare(gi2c->core_clk); + clk_disable_unprepare(gi2c->se.core_clk); return geni_icc_disable(&gi2c->se); } @@ -1214,7 +1212,7 @@ static int __maybe_unused geni_i2c_runtime_resume(struct device *dev) if (ret) return ret; - ret = clk_prepare_enable(gi2c->core_clk); + ret = clk_prepare_enable(gi2c->se.core_clk); if (ret) goto out_icc_disable; @@ -1227,7 +1225,7 @@ static int __maybe_unused geni_i2c_runtime_resume(struct device *dev) return 0; out_clk_disable: - clk_disable_unprepare(gi2c->core_clk); + clk_disable_unprepare(gi2c->se.core_clk); out_icc_disable: geni_icc_disable(&gi2c->se); @@ -1267,15 +1265,24 @@ static const struct dev_pm_ops geni_i2c_pm_ops = { NULL) }; +static const struct geni_i2c_desc geni_i2c = {}; + static const struct geni_i2c_desc i2c_master_hub = { - .has_core_clk = true, - .icc_ddr = NULL, .no_dma_support = true, .tx_fifo_depth = 16, }; +#ifdef CONFIG_ACPI +static const struct acpi_device_id geni_i2c_acpi_match[] = { + { "QCOM0220", (kernel_ulong_t)&geni_i2c}, + { "QCOM0411", (kernel_ulong_t)&geni_i2c}, + { } +}; +MODULE_DEVICE_TABLE(acpi, geni_i2c_acpi_match); +#endif + static const struct of_device_id geni_i2c_dt_match[] = { - { .compatible = "qcom,geni-i2c" }, + { .compatible = "qcom,geni-i2c", .data = &geni_i2c }, { .compatible = "qcom,geni-i2c-master-hub", .data = &i2c_master_hub }, {} }; diff --git a/drivers/net/bonding/bond_main.c b/drivers/net/bonding/bond_main.c index 6ea46a617ee6..34c5bbbed9ae 100644 --- a/drivers/net/bonding/bond_main.c +++ b/drivers/net/bonding/bond_main.c @@ -490,7 +490,7 @@ static int bond_ipsec_add_sa(struct net_device *bond_dev, !real_dev->xfrmdev_ops->xdo_dev_state_add || netif_is_bond_master(real_dev)) { NL_SET_ERR_MSG_MOD(extack, "Slave does not support ipsec offload"); - err = -EINVAL; + err = -EOPNOTSUPP; goto out; } diff --git a/drivers/net/dsa/mt7530-mdio.c b/drivers/net/dsa/mt7530-mdio.c index 784dd58a7158..de42f70afcfa 100644 --- a/drivers/net/dsa/mt7530-mdio.c +++ b/drivers/net/dsa/mt7530-mdio.c @@ -227,15 +227,17 @@ mt7530_remove(struct mdio_device *mdiodev) if (!priv) return; - ret = regulator_disable(priv->core_pwr); - if (ret < 0) - dev_err(priv->dev, - "Failed to disable core power: %d\n", ret); - - ret = regulator_disable(priv->io_pwr); - if (ret < 0) - dev_err(priv->dev, "Failed to disable io pwr: %d\n", - ret); + if (priv->id == ID_MT7530) { + ret = regulator_disable(priv->core_pwr); + if (ret < 0) + dev_err(priv->dev, + "Failed to disable core power: %d\n", ret); + + ret = regulator_disable(priv->io_pwr); + if (ret < 0) + dev_err(priv->dev, "Failed to disable io pwr: %d\n", + ret); + } mt7530_remove_common(priv); diff --git a/drivers/net/dsa/mt7530.c b/drivers/net/dsa/mt7530.c index 56ee8dc34f51..93dba4498c42 100644 --- a/drivers/net/dsa/mt7530.c +++ b/drivers/net/dsa/mt7530.c @@ -3539,9 +3539,6 @@ EXPORT_SYMBOL_GPL(mt7530_probe_common); void mt7530_remove_common(struct mt7530_priv *priv) { - if (priv->irq_domain) - mt7530_free_mdio_irq(priv); - dsa_unregister_switch(priv->ds); mutex_destroy(&priv->reg_mutex); diff --git a/drivers/net/dsa/mv88e6xxx/chip.c b/drivers/net/dsa/mv88e6xxx/chip.c index 7f68a0c55802..a4a8c7e11bf4 100644 --- a/drivers/net/dsa/mv88e6xxx/chip.c +++ b/drivers/net/dsa/mv88e6xxx/chip.c @@ -5639,6 +5639,68 @@ static const struct mv88e6xxx_ops mv88e6390x_ops = { .pcs_ops = &mv88e6390_pcs_ops, }; +static const struct mv88e6xxx_ops mv88e6191x_ops = { + /* MV88E6XXX_FAMILY_6393 without AVB and PTP: 6191X and 6193X */ + .irl_init_all = mv88e6390_g2_irl_init_all, + .get_eeprom = mv88e6xxx_g2_get_eeprom8, + .set_eeprom = mv88e6xxx_g2_set_eeprom8, + .set_switch_mac = mv88e6xxx_g2_set_switch_mac, + .phy_read = mv88e6xxx_g2_smi_phy_read_c22, + .phy_write = mv88e6xxx_g2_smi_phy_write_c22, + .phy_read_c45 = mv88e6xxx_g2_smi_phy_read_c45, + .phy_write_c45 = mv88e6xxx_g2_smi_phy_write_c45, + .port_set_link = mv88e6xxx_port_set_link, + .port_sync_link = mv88e6xxx_port_sync_link, + .port_set_rgmii_delay = mv88e6390_port_set_rgmii_delay, + .port_set_speed_duplex = mv88e6393x_port_set_speed_duplex, + .port_tag_remap = mv88e6390_port_tag_remap, + .port_set_policy = mv88e6393x_port_set_policy, + .port_set_frame_mode = mv88e6351_port_set_frame_mode, + .port_set_ucast_flood = mv88e6352_port_set_ucast_flood, + .port_set_mcast_flood = mv88e6352_port_set_mcast_flood, + .port_set_ether_type = mv88e6393x_port_set_ether_type, + .port_set_jumbo_size = mv88e6165_port_set_jumbo_size, + .port_egress_rate_limiting = mv88e6097_port_egress_rate_limiting, + .port_pause_limit = mv88e6390_port_pause_limit, + .port_disable_learn_limit = mv88e6xxx_port_disable_learn_limit, + .port_disable_pri_override = mv88e6xxx_port_disable_pri_override, + .port_get_cmode = mv88e6352_port_get_cmode, + .port_set_cmode = mv88e6393x_port_set_cmode, + .port_setup_message_port = mv88e6xxx_setup_message_port, + .port_set_upstream_port = mv88e6393x_port_set_upstream_port, + .port_enable_tcam = mv88e6xxx_port_enable_tcam, + .stats_snapshot = mv88e6390_g1_stats_snapshot, + .stats_set_histogram = mv88e6390_g1_stats_set_histogram, + .stats_get_sset_count = mv88e6320_stats_get_sset_count, + .stats_get_strings = mv88e6320_stats_get_strings, + .stats_get_stat = mv88e6390_stats_get_stat, + /* .set_cpu_port is missing because this family does not support a global + * CPU port, only per port CPU port which is set via + * .port_set_upstream_port method. + */ + .set_egress_port = mv88e6393x_set_egress_port, + .watchdog_ops = &mv88e6393x_watchdog_ops, + .mgmt_rsvd2cpu = mv88e6393x_port_mgmt_rsvd2cpu, + .pot_clear = mv88e6xxx_g2_pot_clear, + .hardware_reset_pre = mv88e6xxx_g2_eeprom_wait, + .hardware_reset_post = mv88e6xxx_g2_eeprom_wait, + .reset = mv88e6352_g1_reset, + .rmu_disable = mv88e6390_g1_rmu_disable, + .atu_get_hash = mv88e6165_g1_atu_get_hash, + .atu_set_hash = mv88e6165_g1_atu_set_hash, + .vtu_getnext = mv88e6390_g1_vtu_getnext, + .vtu_loadpurge = mv88e6390_g1_vtu_loadpurge, + .stu_getnext = mv88e6390_g1_stu_getnext, + .stu_loadpurge = mv88e6390_g1_stu_loadpurge, + .serdes_get_lane = mv88e6393x_serdes_get_lane, + .serdes_irq_mapping = mv88e6390_serdes_irq_mapping, + /* TODO: serdes stats */ + .gpio_ops = &mv88e6352_gpio_ops, + .phylink_get_caps = mv88e6393x_phylink_get_caps, + .pcs_ops = &mv88e6393x_pcs_ops, + .tcam_ops = &mv88e6393_tcam_ops, +}; + static const struct mv88e6xxx_ops mv88e6393x_ops = { /* MV88E6XXX_FAMILY_6393 */ .irl_init_all = mv88e6390_g2_irl_init_all, @@ -6163,8 +6225,7 @@ static const struct mv88e6xxx_info mv88e6xxx_table[] = { .atu_move_port_mask = 0x1f, .pvt = true, .multi_chip = true, - .ptp_support = true, - .ops = &mv88e6393x_ops, + .ops = &mv88e6191x_ops, }, [MV88E6193X] = { @@ -6190,8 +6251,7 @@ static const struct mv88e6xxx_info mv88e6xxx_table[] = { .atu_move_port_mask = 0x1f, .pvt = true, .multi_chip = true, - .ptp_support = true, - .ops = &mv88e6393x_ops, + .ops = &mv88e6191x_ops, }, [MV88E6220] = { diff --git a/drivers/net/ethernet/airoha/airoha_npu.c b/drivers/net/ethernet/airoha/airoha_npu.c index 4045d1eb93ea..3862f829f183 100644 --- a/drivers/net/ethernet/airoha/airoha_npu.c +++ b/drivers/net/ethernet/airoha/airoha_npu.c @@ -5,6 +5,7 @@ */ #include +#include #include #include #include @@ -749,12 +750,15 @@ static int airoha_npu_probe(struct platform_device *pdev) if (irq < 0) return irq; + err = devm_work_autocancel(dev, &core->wdt_work, + airoha_npu_wdt_work); + if (err) + return err; + err = devm_request_irq(dev, irq, airoha_npu_wdt_handler, IRQF_SHARED, "airoha-npu-wdt", core); if (err) return err; - - INIT_WORK(&core->wdt_work, airoha_npu_wdt_work); } /* wlan IRQ lines */ @@ -801,18 +805,8 @@ static int airoha_npu_probe(struct platform_device *pdev) return 0; } -static void airoha_npu_remove(struct platform_device *pdev) -{ - struct airoha_npu *npu = platform_get_drvdata(pdev); - int i; - - for (i = 0; i < ARRAY_SIZE(npu->cores); i++) - cancel_work_sync(&npu->cores[i].wdt_work); -} - static struct platform_driver airoha_npu_driver = { .probe = airoha_npu_probe, - .remove = airoha_npu_remove, .driver = { .name = "airoha-npu", .of_match_table = of_airoha_npu_match, diff --git a/drivers/net/ethernet/amazon/ena/ena_netdev.c b/drivers/net/ethernet/amazon/ena/ena_netdev.c index 5d05020a6d05..e35eef9f6ce1 100644 --- a/drivers/net/ethernet/amazon/ena/ena_netdev.c +++ b/drivers/net/ethernet/amazon/ena/ena_netdev.c @@ -4126,6 +4126,8 @@ static int ena_probe(struct pci_dev *pdev, const struct pci_device_id *ent) err_device_destroy: ena_com_delete_host_info(ena_dev); ena_com_admin_destroy(ena_dev); + ena_phc_destroy(adapter); + ena_com_mmio_reg_read_request_destroy(ena_dev); ena_devlink_destroy: ena_devlink_free(devlink); err_metrics_destroy: diff --git a/drivers/net/ethernet/atheros/atl1c/atl1c_main.c b/drivers/net/ethernet/atheros/atl1c/atl1c_main.c index 7efa3fc257b3..e58f1d2c26bd 100644 --- a/drivers/net/ethernet/atheros/atl1c/atl1c_main.c +++ b/drivers/net/ethernet/atheros/atl1c/atl1c_main.c @@ -1602,6 +1602,9 @@ static int atl1c_clean_tx(struct napi_struct *napi, int budget) AT_READ_REGW(&adapter->hw, atl1c_qregs[tpd_ring->num].tpd_cons, &hw_next_to_clean); + if (unlikely(hw_next_to_clean >= tpd_ring->count)) + hw_next_to_clean = next_to_clean; + while (next_to_clean != hw_next_to_clean) { buffer_info = &tpd_ring->buffer_info[next_to_clean]; if (buffer_info->skb) { diff --git a/drivers/net/ethernet/atheros/atl1e/atl1e_main.c b/drivers/net/ethernet/atheros/atl1e/atl1e_main.c index 40290028580b..437989edb780 100644 --- a/drivers/net/ethernet/atheros/atl1e/atl1e_main.c +++ b/drivers/net/ethernet/atheros/atl1e/atl1e_main.c @@ -1234,6 +1234,9 @@ static bool atl1e_clean_tx_irq(struct atl1e_adapter *adapter) u16 hw_next_to_clean = AT_READ_REGW(&adapter->hw, REG_TPD_CONS_IDX); u16 next_to_clean = atomic_read(&tx_ring->next_to_clean); + if (unlikely(hw_next_to_clean >= tx_ring->count)) + hw_next_to_clean = next_to_clean; + while (next_to_clean != hw_next_to_clean) { tx_buffer = &tx_ring->tx_buffer[next_to_clean]; if (tx_buffer->dma) { diff --git a/drivers/net/ethernet/atheros/atlx/atl1.c b/drivers/net/ethernet/atheros/atlx/atl1.c index 98a4d089270e..957d5598dda5 100644 --- a/drivers/net/ethernet/atheros/atlx/atl1.c +++ b/drivers/net/ethernet/atheros/atlx/atl1.c @@ -2066,6 +2066,9 @@ static int atl1_intr_tx(struct atl1_adapter *adapter) sw_tpd_next_to_clean = atomic_read(&tpd_ring->next_to_clean); cmb_tpd_next_to_clean = le16_to_cpu(adapter->cmb.cmb->tpd_cons_idx); + if (unlikely(cmb_tpd_next_to_clean >= tpd_ring->count)) + cmb_tpd_next_to_clean = sw_tpd_next_to_clean; + while (cmb_tpd_next_to_clean != sw_tpd_next_to_clean) { buffer_info = &tpd_ring->buffer_info[sw_tpd_next_to_clean]; if (buffer_info->dma) { diff --git a/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c b/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c index 62bc9cae613c..622e89587e5d 100644 --- a/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c +++ b/drivers/net/ethernet/broadcom/bnxt/bnxt_ethtool.c @@ -5733,7 +5733,8 @@ const struct ethtool_ops bnxt_ethtool_ops = { .op_needs_rtnl = ETHTOOL_OP_NEEDS_RTNL_SCHANNELS | ETHTOOL_OP_NEEDS_RTNL_SRINGPARAM | ETHTOOL_OP_NEEDS_RTNL_SCOALESCE | - ETHTOOL_OP_NEEDS_RTNL_RSS, + ETHTOOL_OP_NEEDS_RTNL_RSS | + ETHTOOL_OP_NEEDS_RTNL_TEST, .supported_coalesce_params = ETHTOOL_COALESCE_USECS | ETHTOOL_COALESCE_MAX_FRAMES | ETHTOOL_COALESCE_USECS_IRQ | diff --git a/drivers/net/ethernet/broadcom/genet/bcmgenet.c b/drivers/net/ethernet/broadcom/genet/bcmgenet.c index b916080f4ff1..21668e41b696 100644 --- a/drivers/net/ethernet/broadcom/genet/bcmgenet.c +++ b/drivers/net/ethernet/broadcom/genet/bcmgenet.c @@ -852,7 +852,8 @@ static int bcmgenet_get_coalesce(struct net_device *dev, ec->rx_max_coalesced_frames = bcmgenet_rdma_ring_readl(priv, 0, DMA_MBUF_DONE_THRESH); ec->rx_coalesce_usecs = - bcmgenet_rdma_readl(priv, DMA_RING0_TIMEOUT) * 8192 / 1000; + (bcmgenet_rdma_readl(priv, DMA_RING0_TIMEOUT) & + DMA_TIMEOUT_MASK) * 8192 / 1000; for (i = 0; i <= priv->hw_params->rx_queues; i++) { ring = &priv->rx_rings[i]; @@ -1346,9 +1347,8 @@ static void bcmgenet_get_ethtool_stats(struct net_device *dev, p = (char *)&stats64; p += s->stat_offset; - if (sizeof(unsigned long) != sizeof(u32) && - s->stat_sizeof == sizeof(unsigned long)) - data[i] = *(unsigned long *)p; + if (s->stat_sizeof == sizeof(u64)) + data[i] = *(u64 *)p; else data[i] = *(u32 *)p; } @@ -1763,13 +1763,12 @@ static int bcmgenet_power_up(struct bcmgenet_priv *priv, int ret = 0; u32 reg; - if (!bcmgenet_has_ext(priv)) - return ret; - - reg = bcmgenet_ext_readl(priv, EXT_EXT_PWR_MGMT); - switch (mode) { case GENET_POWER_PASSIVE: + if (!bcmgenet_has_ext(priv)) + break; + + reg = bcmgenet_ext_readl(priv, EXT_EXT_PWR_MGMT); reg &= ~(EXT_PWR_DOWN_DLL | EXT_PWR_DOWN_BIAS | EXT_ENERGY_DET_MASK); if (GENET_IS_V5(priv) && !bcmgenet_has_ephy_16nm(priv)) { @@ -1793,8 +1792,12 @@ static int bcmgenet_power_up(struct bcmgenet_priv *priv, break; case GENET_POWER_CABLE_SENSE: + if (!bcmgenet_has_ext(priv)) + break; + /* enable APD */ if (!GENET_IS_V5(priv)) { + reg = bcmgenet_ext_readl(priv, EXT_EXT_PWR_MGMT); reg |= EXT_PWR_DN_EN_LD; bcmgenet_ext_writel(priv, reg, EXT_EXT_PWR_MGMT); } @@ -3441,6 +3444,8 @@ static void bcmgenet_netif_stop(struct net_device *dev, bool stop_phy) { struct bcmgenet_priv *priv = netdev_priv(dev); + /* Stop completion polling before it can wake a stopped queue */ + bcmgenet_disable_tx_napi(priv); netif_tx_disable(dev); /* Disable MAC receive */ @@ -3455,7 +3460,6 @@ static void bcmgenet_netif_stop(struct net_device *dev, bool stop_phy) /* Disable MAC transmit. TX DMA disabled must be done before this */ umac_enable_set(priv, CMD_TX_EN, false); - bcmgenet_disable_tx_napi(priv); bcmgenet_disable_rx_napi(priv); bcmgenet_intr_disable(priv); @@ -3632,6 +3636,9 @@ static int bcmgenet_set_mac_addr(struct net_device *dev, void *p) if (netif_running(dev)) return -EBUSY; + if (!is_valid_ether_addr(addr->sa_data)) + return -EADDRNOTAVAIL; + eth_hw_addr_set(dev, addr->sa_data); return 0; @@ -4135,10 +4142,10 @@ static int bcmgenet_probe(struct platform_device *pdev) priv->rx_rings[i].rx_max_coalesced_frames = 1; /* Initialize u64 stats seq counter for 32bit machines */ - for (i = 0; i <= priv->hw_params->rx_queues; i++) + for (i = 0; i <= GENET_MAX_MQ_CNT; i++) { u64_stats_init(&priv->rx_rings[i].stats64.syncp); - for (i = 0; i <= priv->hw_params->tx_queues; i++) u64_stats_init(&priv->tx_rings[i].stats64.syncp); + } /* libphy will determine the link state */ netif_carrier_off(dev); @@ -4320,6 +4327,8 @@ static int bcmgenet_suspend(struct device *d) netif_device_detach(dev); if (device_may_wakeup(d) && priv->wolopts) { + /* Stop completion polling before it can wake a stopped queue */ + bcmgenet_disable_tx_napi(priv); netif_tx_disable(dev); /* Suspend non-wake Rx data flows */ @@ -4348,7 +4357,6 @@ static int bcmgenet_suspend(struct device *d) netdev_warn(priv->dev, "Timed out while disabling TX DMA\n"); - bcmgenet_disable_tx_napi(priv); bcmgenet_disable_rx_napi(priv); disable_irq(priv->irq1); bcmgenet_tx_reclaim_all(dev); diff --git a/drivers/net/ethernet/broadcom/tg3.c b/drivers/net/ethernet/broadcom/tg3.c index 73a4b569b03e..8b6806a79edf 100644 --- a/drivers/net/ethernet/broadcom/tg3.c +++ b/drivers/net/ethernet/broadcom/tg3.c @@ -17915,11 +17915,14 @@ static int tg3_init_one(struct pci_dev *pdev, err = tg3_get_device_address(tp, addr); if (err) { - dev_err(&pdev->dev, - "Could not obtain valid ethernet address, aborting\n"); - goto err_out_apeunmap; + dev_warn_probe(&pdev->dev, err, + "Could not obtain a valid ethernet address\n"); + if (err == -EPROBE_DEFER) + goto err_out_apeunmap; + eth_hw_addr_random(dev); + } else { + eth_hw_addr_set(dev, addr); } - eth_hw_addr_set(dev, addr); intmbx = MAILBOX_INTERRUPT_0 + TG3_64BIT_REG_LOW; rcvmbx = MAILBOX_RCVRET_CON_IDX_0 + TG3_64BIT_REG_LOW; @@ -18047,6 +18050,10 @@ static int tg3_init_one(struct pci_dev *pdev, return 0; err_out_apeunmap: + if (tg3_flag(tp, USE_PHYLIB)) + tg3_phy_fini(tp); + tg3_mdio_fini(tp); + if (tp->aperegs) { iounmap(tp->aperegs); tp->aperegs = NULL; diff --git a/drivers/net/ethernet/brocade/bna/bnad.c b/drivers/net/ethernet/brocade/bna/bnad.c index 8e19add764db..3fa805117dde 100644 --- a/drivers/net/ethernet/brocade/bna/bnad.c +++ b/drivers/net/ethernet/brocade/bna/bnad.c @@ -2571,6 +2571,22 @@ bnad_ioceth_disable(struct bnad *bnad) return err; } +/* + * The IOC timers rearm one another, so deleting one cannot stop a + * sibling callback from arming it again. Shut them down so a later + * mod_timer() is ignored. + */ +static void +bnad_ioc_timers_shutdown(struct bnad *bnad) +{ + struct bfa_ioc *ioc = &bnad->bna.ioceth.ioc; + + timer_shutdown_sync(&ioc->ioc_timer); + timer_shutdown_sync(&ioc->sem_timer); + timer_shutdown_sync(&ioc->hb_timer); + timer_shutdown_sync(&ioc->iocpf_timer); +} + static int bnad_ioceth_enable(struct bnad *bnad) { @@ -3727,9 +3743,7 @@ bnad_pci_probe(struct pci_dev *pdev, bnad_res_free(bnad, &bnad->mod_res_info[0], BNA_MOD_RES_T_MAX); disable_ioceth: bnad_ioceth_disable(bnad); - timer_delete_sync(&bnad->bna.ioceth.ioc.ioc_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.sem_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.hb_timer); + bnad_ioc_timers_shutdown(bnad); spin_lock_irqsave(&bnad->bna_lock, flags); bna_uninit(bna); spin_unlock_irqrestore(&bnad->bna_lock, flags); @@ -3770,9 +3784,7 @@ bnad_pci_remove(struct pci_dev *pdev) mutex_lock(&bnad->conf_mutex); bnad_ioceth_disable(bnad); - timer_delete_sync(&bnad->bna.ioceth.ioc.ioc_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.sem_timer); - timer_delete_sync(&bnad->bna.ioceth.ioc.hb_timer); + bnad_ioc_timers_shutdown(bnad); spin_lock_irqsave(&bnad->bna_lock, flags); bna_uninit(bna); spin_unlock_irqrestore(&bnad->bna_lock, flags); diff --git a/drivers/net/ethernet/cadence/macb_main.c b/drivers/net/ethernet/cadence/macb_main.c index 8085a3b3846d..a8d830eeb8ea 100644 --- a/drivers/net/ethernet/cadence/macb_main.c +++ b/drivers/net/ethernet/cadence/macb_main.c @@ -2756,14 +2756,24 @@ static int macb_alloc_consistent(struct macb *bp) size = bp->num_queues * macb_tx_ring_size_per_queue(bp); tx = dma_alloc_coherent(dev, size, &tx_dma, GFP_KERNEL); - if (!tx || upper_32_bits(tx_dma) != upper_32_bits(tx_dma + size - 1)) + if (!tx) + goto out_err; + /* Record the buffer so that the error path frees it. */ + bp->queues[0].tx_ring = tx; + bp->queues[0].tx_ring_dma = tx_dma; + if (upper_32_bits(tx_dma) != upper_32_bits(tx_dma + size - 1)) goto out_err; netdev_dbg(bp->netdev, "Allocated %zu bytes for %u TX rings at %08lx (mapped %p)\n", size, bp->num_queues, (unsigned long)tx_dma, tx); size = bp->num_queues * macb_rx_ring_size_per_queue(bp); rx = dma_alloc_coherent(dev, size, &rx_dma, GFP_KERNEL); - if (!rx || upper_32_bits(rx_dma) != upper_32_bits(rx_dma + size - 1)) + if (!rx) + goto out_err; + /* Record the buffer so that the error path frees it. */ + bp->queues[0].rx_ring = rx; + bp->queues[0].rx_ring_dma = rx_dma; + if (upper_32_bits(rx_dma) != upper_32_bits(rx_dma + size - 1)) goto out_err; netdev_dbg(bp->netdev, "Allocated %zu bytes for %u RX rings at %08lx (mapped %p)\n", size, bp->num_queues, (unsigned long)rx_dma, rx); diff --git a/drivers/net/ethernet/freescale/fman/fman.c b/drivers/net/ethernet/freescale/fman/fman.c index 299bab043175..46cc28895e56 100644 --- a/drivers/net/ethernet/freescale/fman/fman.c +++ b/drivers/net/ethernet/freescale/fman/fman.c @@ -2734,6 +2734,7 @@ static struct fman *read_dts_node(struct platform_device *of_dev) } clk_rate = clk_get_rate(clk); + clk_put(clk); if (!clk_rate) { err = -EINVAL; dev_err(&of_dev->dev, "%s: Failed to determine FM%d clock rate\n", diff --git a/drivers/net/ethernet/google/gve/gve_desc_dqo.h b/drivers/net/ethernet/google/gve/gve_desc_dqo.h index f7786b03c744..d2c86c8eeae2 100644 --- a/drivers/net/ethernet/google/gve/gve_desc_dqo.h +++ b/drivers/net/ethernet/google/gve/gve_desc_dqo.h @@ -14,6 +14,11 @@ #define GVE_TX_MAX_HDR_SIZE_DQO 255 #define GVE_TX_MIN_TSO_MSS_DQO 88 +/* HW limit. This also has to fit in the 14 bits of the mss field of + * struct gve_tx_tso_context_desc_dqo. + */ +#define GVE_TX_MAX_TSO_MSS_DQO 9728 + #ifndef __LITTLE_ENDIAN_BITFIELD #error "Only little endian supported" #endif diff --git a/drivers/net/ethernet/google/gve/gve_tx_dqo.c b/drivers/net/ethernet/google/gve/gve_tx_dqo.c index 80ab0a449ff5..616c1921aebe 100644 --- a/drivers/net/ethernet/google/gve/gve_tx_dqo.c +++ b/drivers/net/ethernet/google/gve/gve_tx_dqo.c @@ -577,15 +577,18 @@ static int gve_prep_tso(struct sk_buff *skb) int header_len; int err; - /* Note: HW requires MSS (gso_size) to be <= 9728 and the total length - * of the TSO to be <= 262143. + /* Note: HW requires the total length of the TSO to be <= 262143, + * this is enforced by netif_set_tso_max_size(). * - * However, we don't validate these because: - * - Hypervisor enforces a limit of 9K MTU - * - Kernel will not produce a TSO larger than 64k + * MSS (gso_size) can not be trusted: packets forwarded from a tap or + * injected by a packet socket can carry an arbitrary value, while the + * mss field of the TSO context descriptor is only 14 bits wide. + * + * A too big MSS is dropped here instead of being rejected from + * gve_features_check_dqo(), because software segmentation would + * produce packets larger than the device can send. */ - - if (unlikely(shinfo->gso_size < GVE_TX_MIN_TSO_MSS_DQO)) + if (unlikely(shinfo->gso_size > GVE_TX_MAX_TSO_MSS_DQO)) return -1; /* Needed because we will modify header. */ @@ -918,13 +921,22 @@ static bool gve_can_send_tso(const struct sk_buff *skb) { const int max_bufs_per_seg = GVE_TX_MAX_DATA_DESCS - 1; const struct skb_shared_info *shinfo = skb_shinfo(skb); - const int header_len = skb_tcp_all_headers(skb); const int gso_size = shinfo->gso_size; int cur_seg_num_bufs; int prev_frag_size; int cur_seg_size; + int header_len; int i; + if (unlikely(gso_size < GVE_TX_MIN_TSO_MSS_DQO)) + return false; + + /* Must match the header length programmed by gve_prep_tso(). */ + if (skb_is_gso_tcp(skb)) + header_len = skb_tcp_all_headers(skb); + else + header_len = skb_transport_offset(skb) + sizeof(struct udphdr); + cur_seg_size = skb_headlen(skb) - header_len; prev_frag_size = skb_headlen(skb); cur_seg_num_bufs = cur_seg_size > 0; @@ -966,7 +978,17 @@ netdev_features_t gve_features_check_dqo(struct sk_buff *skb, struct net_device *dev, netdev_features_t features) { - if (skb_is_gso(skb) && !gve_can_send_tso(skb)) + if (!skb_is_gso(skb)) + return features; + + /* Keep the GSO bits for a too big MSS, so that gve_prep_tso() drops + * the packet: software segmentation would give packets larger than + * the device can send. + */ + if (skb_shinfo(skb)->gso_size > GVE_TX_MAX_TSO_MSS_DQO) + return features; + + if (!gve_can_send_tso(skb)) return features & ~NETIF_F_GSO_MASK; return features; diff --git a/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c b/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c index bc6b269be299..f1cb6d56e4b9 100644 --- a/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c +++ b/drivers/net/ethernet/hisilicon/hns/hns_dsaf_mac.c @@ -793,6 +793,7 @@ static int hns_mac_register_phy(struct hns_mac_cb *mac_cb) dev_err(mac_cb->dev, "mac%d mdio is NULL, dsaf will probe again later\n", mac_cb->mac_id); + put_device(&pdev->dev); return -EPROBE_DEFER; } @@ -801,6 +802,8 @@ static int hns_mac_register_phy(struct hns_mac_cb *mac_cb) dev_dbg(mac_cb->dev, "mac%d register phy addr:%d\n", mac_cb->mac_id, addr); + put_device(&pdev->dev); + return rc; } diff --git a/drivers/net/ethernet/ibm/emac/core.c b/drivers/net/ethernet/ibm/emac/core.c index 1d46cf6c2c12..e7043523457c 100644 --- a/drivers/net/ethernet/ibm/emac/core.c +++ b/drivers/net/ethernet/ibm/emac/core.c @@ -3044,6 +3044,15 @@ static int emac_probe(struct platform_device *ofdev) if (err) goto err_gone; + if (emac_phy_supports_gige(dev->phy_mode)) { + ndev->netdev_ops = &emac_gige_netdev_ops; + dev->commac.ops = &emac_commac_sg_ops; + } else { + ndev->netdev_ops = &emac_netdev_ops; + dev->commac.ops = &emac_commac_ops; + } + ndev->ethtool_ops = &emac_ethtool_ops; + dev->emacp = devm_platform_ioremap_resource(ofdev, 0); if (IS_ERR(dev->emacp)) { err = PTR_ERR(dev->emacp); @@ -3076,7 +3085,6 @@ static int emac_probe(struct platform_device *ofdev) dev->mdio_instance = platform_get_drvdata(dev->mdio_dev); /* Register with MAL */ - dev->commac.ops = &emac_commac_ops; dev->commac.dev = dev; dev->commac.tx_chan_mask = MAL_CHAN_MASK(dev->mal_tx_chan); dev->commac.rx_chan_mask = MAL_CHAN_MASK(dev->mal_rx_chan); @@ -3144,12 +3152,6 @@ static int emac_probe(struct platform_device *ofdev) ndev->features |= ndev->hw_features | NETIF_F_RXCSUM; } ndev->watchdog_timeo = 5 * HZ; - if (emac_phy_supports_gige(dev->phy_mode)) { - ndev->netdev_ops = &emac_gige_netdev_ops; - dev->commac.ops = &emac_commac_sg_ops; - } else - ndev->netdev_ops = &emac_netdev_ops; - ndev->ethtool_ops = &emac_ethtool_ops; /* MTU range: 46 - 1500 or whatever is in OF */ ndev->min_mtu = EMAC_MIN_MTU; diff --git a/drivers/net/ethernet/marvell/octeontx2/af/common.h b/drivers/net/ethernet/marvell/octeontx2/af/common.h index 779413a383b7..78e42549d990 100644 --- a/drivers/net/ethernet/marvell/octeontx2/af/common.h +++ b/drivers/net/ethernet/marvell/octeontx2/af/common.h @@ -7,6 +7,10 @@ #ifndef COMMON_H #define COMMON_H +#include +#include +#include + #include "rvu_struct.h" #define OTX2_ALIGN 128 /* Align to cacheline */ @@ -44,6 +48,33 @@ struct qmem { u32 qsize; }; +static inline void *otx2_dma_alloc_coherent(struct device *dev, size_t size, + dma_addr_t *dma_handle) +{ + dma_addr_t dma_addr; + void *vaddr; + + vaddr = kzalloc(size, GFP_KERNEL); + if (!vaddr) + return NULL; + + dma_addr = dma_map_single(dev, vaddr, size, DMA_BIDIRECTIONAL); + if (dma_mapping_error(dev, dma_addr)) { + kfree(vaddr); + return NULL; + } + + *dma_handle = dma_addr; + return vaddr; +} + +static inline void otx2_dma_free_coherent(struct device *dev, size_t size, + void *vaddr, dma_addr_t dma_handle) +{ + dma_unmap_single(dev, dma_handle, size, DMA_BIDIRECTIONAL); + kfree(vaddr); +} + static inline int qmem_alloc(struct device *dev, struct qmem **q, int qsize, int entry_sz) { @@ -60,8 +91,11 @@ static inline int qmem_alloc(struct device *dev, struct qmem **q, qmem->entry_sz = entry_sz; qmem->alloc_sz = (qsize * entry_sz) + OTX2_ALIGN; - qmem->base = dma_alloc_attrs(dev, qmem->alloc_sz, &qmem->iova, - GFP_KERNEL, DMA_ATTR_FORCE_CONTIGUOUS); + + if (get_order(PAGE_ALIGN(qmem->alloc_sz)) > MAX_PAGE_ORDER) + return -ENOMEM; + + qmem->base = otx2_dma_alloc_coherent(dev, qmem->alloc_sz, &qmem->iova); if (!qmem->base) return -ENOMEM; @@ -80,10 +114,9 @@ static inline void qmem_free(struct device *dev, struct qmem *qmem) return; if (qmem->base) - dma_free_attrs(dev, qmem->alloc_sz, - qmem->base - qmem->align, - qmem->iova - qmem->align, - DMA_ATTR_FORCE_CONTIGUOUS); + otx2_dma_free_coherent(dev, qmem->alloc_sz, + qmem->base - qmem->align, + qmem->iova - qmem->align); devm_kfree(dev, qmem); } diff --git a/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c b/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c index 904374baae6f..2927633465d9 100644 --- a/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c +++ b/drivers/net/ethernet/marvell/octeontx2/af/rvu_debugfs.c @@ -714,110 +714,72 @@ static int get_max_column_width(struct rvu *rvu) } /* Dumps current provisioning status of all RVU block LFs */ -static ssize_t rvu_dbg_rsrc_attach_status(struct file *filp, - char __user *buffer, - size_t count, loff_t *ppos) +static int rvu_dbg_rsrc_attach_status(struct seq_file *filp, void *unused) { - int index, off = 0, flag = 0, len = 0, i = 0; - struct rvu *rvu = filp->private_data; - int bytes_not_copied = 0; + struct rvu *rvu = filp->private; + int index, pf, vf, pcifunc; struct rvu_block block; - int pf, vf, pcifunc; - int buf_size = 2048; int lf_str_size; char *lfs; - char *buf; - - /* don't allow partial reads */ - if (*ppos != 0) - return 0; - - buf = kzalloc(buf_size, GFP_KERNEL); - if (!buf) - return -ENOMEM; - /* Get the maximum width of a column */ lf_str_size = get_max_column_width(rvu); + if (lf_str_size < 0) + return lf_str_size; lfs = kzalloc(lf_str_size, GFP_KERNEL); - if (!lfs) { - kfree(buf); + if (!lfs) return -ENOMEM; - } - off += scnprintf(&buf[off], buf_size - 1 - off, "%-*s", lf_str_size, - "pcifunc"); - for (index = 0; index < BLK_COUNT; index++) - if (strlen(rvu->hw->block[index].name)) { - off += scnprintf(&buf[off], buf_size - 1 - off, - "%-*s", lf_str_size, - rvu->hw->block[index].name); - } - off += scnprintf(&buf[off], buf_size - 1 - off, "\n"); - bytes_not_copied = copy_to_user(buffer + (i * off), buf, off); - if (bytes_not_copied) - goto out; + seq_printf(filp, "%-*s", lf_str_size, "pcifunc"); + for (index = 0; index < BLK_COUNT; index++) + if (strlen(rvu->hw->block[index].name)) + seq_printf(filp, "%-*s", lf_str_size, + rvu->hw->block[index].name); - i++; - *ppos += off; + seq_putc(filp, '\n'); for (pf = 0; pf < rvu->hw->total_pfs; pf++) { for (vf = 0; vf <= rvu->hw->total_vfs; vf++) { - off = 0; - flag = 0; pcifunc = rvu_make_pcifunc(rvu->pdev, pf, vf); if (!pcifunc) continue; - if (vf) { + for (index = 0; index < BLK_COUNT; index++) { + block = rvu->hw->block[index]; + if (!strlen(block.name)) + continue; + lfs[0] = '\0'; + get_lf_str_list(&block, pcifunc, lfs); + if (strlen(lfs)) + break; + } + if (index == BLK_COUNT) + continue; + + if (vf) sprintf(lfs, "PF%d:VF%d", pf, vf - 1); - off = scnprintf(&buf[off], - buf_size - 1 - off, - "%-*s", lf_str_size, lfs); - } else { + else sprintf(lfs, "PF%d", pf); - off = scnprintf(&buf[off], - buf_size - 1 - off, - "%-*s", lf_str_size, lfs); - } + seq_printf(filp, "%-*s", lf_str_size, lfs); for (index = 0; index < BLK_COUNT; index++) { block = rvu->hw->block[index]; if (!strlen(block.name)) continue; - len = 0; - lfs[len] = '\0'; - get_lf_str_list(&block, pcifunc, lfs); - if (strlen(lfs)) - flag = 1; - off += scnprintf(&buf[off], buf_size - 1 - off, - "%-*s", lf_str_size, lfs); - } - if (flag) { - off += scnprintf(&buf[off], - buf_size - 1 - off, "\n"); - bytes_not_copied = copy_to_user(buffer + - (i * off), - buf, off); - if (bytes_not_copied) - goto out; - - i++; - *ppos += off; + lfs[0] = '\0'; + get_lf_str_list(&block, pcifunc, lfs); + seq_printf(filp, "%-*s", lf_str_size, lfs); } + seq_putc(filp, '\n'); } } -out: kfree(lfs); - kfree(buf); - if (bytes_not_copied) - return -EFAULT; - return *ppos; + return 0; } -RVU_DEBUG_FOPS(rsrc_status, rsrc_attach_status, NULL); +RVU_DEBUG_SEQ_FOPS(rsrc_status, rsrc_attach_status, NULL); static int rvu_dbg_rvu_pf_cgx_map_display(struct seq_file *filp, void *unused) { diff --git a/drivers/net/ethernet/mediatek/mtk_eth_soc.c b/drivers/net/ethernet/mediatek/mtk_eth_soc.c index b2473df74dff..efc238abb67b 100644 --- a/drivers/net/ethernet/mediatek/mtk_eth_soc.c +++ b/drivers/net/ethernet/mediatek/mtk_eth_soc.c @@ -4508,6 +4508,10 @@ static int mtk_unreg_dev(struct mtk_eth *eth) mac = netdev_priv(eth->netdev[i]); if (MTK_HAS_CAPS(eth->soc->caps, MTK_QDMA)) unregister_netdevice_notifier(&mac->device_notifier); + + if (eth->netdev[i]->reg_state != NETREG_REGISTERED) + continue; + unregister_netdev(eth->netdev[i]); } @@ -5339,7 +5343,7 @@ static int mtk_probe(struct platform_device *pdev) err = register_netdev(eth->netdev[i]); if (err) { dev_err(eth->dev, "error bringing up device\n"); - goto err_deinit_ppe; + goto err_unreg_netdev; } else netif_info(eth, probe, eth->netdev[i], "mediatek frame engine at 0x%08lx, irq %d\n", diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c b/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c index baac38bece14..4b7b0a0fc2b2 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en/rep/bridge.c @@ -85,9 +85,16 @@ mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(struct net_device *dev, struct m struct net_device *lower_dev; struct list_head *iter; - if (netif_is_lag_master(dev) || mlx5e_eswitch_rep(dev)) - return mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, esw, vport_num, - esw_owner_vhca_id); + if (netif_is_lag_master(dev) || mlx5e_eswitch_rep(dev)) { + struct net_device *rep; + + rep = mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, esw, vport_num, + esw_owner_vhca_id); + if (rep && !mlx5_esw_bridge_port_exists(*vport_num, *esw_owner_vhca_id, + esw->br_offloads)) + return NULL; + return rep; + } netdev_for_each_lower_dev(dev, lower_dev, iter) { struct net_device *rep; @@ -104,6 +111,28 @@ mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(struct net_device *dev, struct m return NULL; } +static bool mlx5_esw_bridge_rep_port_lookup(struct net_device *dev, + struct mlx5_esw_bridge_offloads *br_offloads, + u16 *vport_num, u16 *esw_owner_vhca_id) +{ + if (!mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, br_offloads->esw, vport_num, + esw_owner_vhca_id)) + return false; + + return mlx5_esw_bridge_port_exists(*vport_num, *esw_owner_vhca_id, br_offloads); +} + +static bool mlx5_esw_bridge_lower_rep_port_lookup(struct net_device *dev, + struct mlx5_esw_bridge_offloads *br_offloads, + u16 *vport_num, u16 *esw_owner_vhca_id) +{ + if (!mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(dev, br_offloads->esw, vport_num, + esw_owner_vhca_id)) + return false; + + return mlx5_esw_bridge_port_exists(*vport_num, *esw_owner_vhca_id, br_offloads); +} + static bool mlx5_esw_bridge_is_local(struct net_device *dev, struct net_device *rep, struct mlx5_eswitch *esw) { @@ -218,8 +247,7 @@ mlx5_esw_bridge_port_obj_add(struct net_device *dev, u16 vport_num, esw_owner_vhca_id; int err; - if (!mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, br_offloads->esw, &vport_num, - &esw_owner_vhca_id)) + if (!mlx5_esw_bridge_rep_port_lookup(dev, br_offloads, &vport_num, &esw_owner_vhca_id)) return 0; port_obj_info->handled = true; @@ -251,8 +279,7 @@ mlx5_esw_bridge_port_obj_del(struct net_device *dev, const struct switchdev_obj_port_mdb *mdb; u16 vport_num, esw_owner_vhca_id; - if (!mlx5_esw_bridge_rep_vport_num_vhca_id_get(dev, br_offloads->esw, &vport_num, - &esw_owner_vhca_id)) + if (!mlx5_esw_bridge_rep_port_lookup(dev, br_offloads, &vport_num, &esw_owner_vhca_id)) return 0; port_obj_info->handled = true; @@ -283,8 +310,8 @@ mlx5_esw_bridge_port_obj_attr_set(struct net_device *dev, u16 vport_num, esw_owner_vhca_id; int err = 0; - if (!mlx5_esw_bridge_lower_rep_vport_num_vhca_id_get(dev, br_offloads->esw, &vport_num, - &esw_owner_vhca_id)) + if (!mlx5_esw_bridge_lower_rep_port_lookup(dev, br_offloads, &vport_num, + &esw_owner_vhca_id)) return 0; port_attr_info->handled = true; diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c b/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c index 6c87a1c7db09..c1841ce74d9a 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en/tc_ct.c @@ -1082,6 +1082,9 @@ mlx5_tc_ct_shared_counter_get(struct mlx5_tc_ct_priv *ct_priv, spin_unlock_bh(&ct_priv->ht_lock); + if (rev_entry) + mlx5_tc_ct_entry_put(rev_entry); + create_counter: shared_counter = mlx5_tc_ct_counter_create(ct_priv); diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c index 329608c59313..8ffa8068e90a 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/ipsec_fs.c @@ -1564,14 +1564,14 @@ static void setup_fte_addr6(struct mlx5_flow_spec *spec, memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_value, outer_headers.src_ipv4_src_ipv6.ipv6_layout.ipv6), saddr, 16); memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_criteria, - outer_headers.src_ipv4_src_ipv6.ipv6_layout.ipv6), dmask, 16); + outer_headers.src_ipv4_src_ipv6.ipv6_layout.ipv6), smask, 16); } if (!addr6_all_zero(daddr)) { memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_value, outer_headers.dst_ipv4_dst_ipv6.ipv6_layout.ipv6), daddr, 16); memcpy(MLX5_ADDR_OF(fte_match_param, spec->match_criteria, - outer_headers.dst_ipv4_dst_ipv6.ipv6_layout.ipv6), smask, 16); + outer_headers.dst_ipv4_dst_ipv6.ipv6_layout.ipv6), dmask, 16); } } diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c index daff53ba7d09..38a3415acf7a 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_accel/macsec.c @@ -1724,6 +1724,8 @@ void mlx5e_macsec_build_netdev(struct mlx5e_priv *priv) mlx5_core_dbg(priv->mdev, "mlx5e: MACsec acceleration enabled\n"); netdev->macsec_ops = &macsec_offload_ops; netdev->features |= NETIF_F_HW_MACSEC; + netdev->hw_features |= NETIF_F_HW_MACSEC; + netdev->vlan_features |= NETIF_F_HW_MACSEC; netif_keep_dst(netdev); } diff --git a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c index 8634c970cd00..da907248dc6c 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/en_main.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/en_main.c @@ -5851,7 +5851,6 @@ static void mlx5e_build_nic_netdev(struct net_device *netdev) netdev->vlan_features |= NETIF_F_SG; netdev->vlan_features |= NETIF_F_HW_CSUM; - netdev->vlan_features |= NETIF_F_HW_MACSEC; netdev->vlan_features |= NETIF_F_GRO; netdev->vlan_features |= NETIF_F_TSO; netdev->vlan_features |= NETIF_F_TSO6; diff --git a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c index 87b5fd349594..b4cf3c5ac0dd 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.c @@ -1649,10 +1649,8 @@ int mlx5_esw_bridge_vport_unlink(struct net_device *br_netdev, u16 vport_num, int err; port = mlx5_esw_bridge_port_lookup(vport_num, esw_owner_vhca_id, br_offloads); - if (!port) { - NL_SET_ERR_MSG_MOD(extack, "Port is not attached to any bridge"); - return -EINVAL; - } + if (!port) + return 0; if (port->bridge->ifindex != br_netdev->ifindex) { NL_SET_ERR_MSG_MOD(extack, "Port is attached to another bridge"); return -EINVAL; @@ -1682,10 +1680,19 @@ int mlx5_esw_bridge_vport_peer_unlink(struct net_device *br_netdev, u16 vport_nu struct mlx5_esw_bridge_offloads *br_offloads, struct netlink_ext_ack *extack) { + if (!MLX5_CAP_ESW(br_offloads->esw->dev, merged_eswitch)) + return 0; + return mlx5_esw_bridge_vport_unlink(br_netdev, vport_num, esw_owner_vhca_id, br_offloads, extack); } +bool mlx5_esw_bridge_port_exists(u16 vport_num, u16 esw_owner_vhca_id, + struct mlx5_esw_bridge_offloads *br_offloads) +{ + return mlx5_esw_bridge_port_lookup(vport_num, esw_owner_vhca_id, br_offloads); +} + int mlx5_esw_bridge_port_vlan_add(u16 vport_num, u16 esw_owner_vhca_id, u16 vid, u16 flags, struct mlx5_esw_bridge_offloads *br_offloads, struct netlink_ext_ack *extack) diff --git a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h index d6f539161993..a4e59cc21089 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h +++ b/drivers/net/ethernet/mellanox/mlx5/core/esw/bridge.h @@ -80,6 +80,8 @@ int mlx5_esw_bridge_vlan_proto_set(u16 vport_num, u16 esw_owner_vhca_id, u16 pro struct mlx5_esw_bridge_offloads *br_offloads); int mlx5_esw_bridge_mcast_set(u16 vport_num, u16 esw_owner_vhca_id, bool enable, struct mlx5_esw_bridge_offloads *br_offloads); +bool mlx5_esw_bridge_port_exists(u16 vport_num, u16 esw_owner_vhca_id, + struct mlx5_esw_bridge_offloads *br_offloads); int mlx5_esw_bridge_port_vlan_add(u16 vport_num, u16 esw_owner_vhca_id, u16 vid, u16 flags, struct mlx5_esw_bridge_offloads *br_offloads, struct netlink_ext_ack *extack); diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c index c655f6e32e9b..dd14cdc378de 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/lag.c @@ -1266,25 +1266,45 @@ void mlx5_lag_remove_devices(struct mlx5_lag *ldev) mlx5_lag_remove_devices_filter(ldev, MLX5_LAG_FILTER_PORTS); } +static int mlx5_lag_reload_ib_reps_idx(struct mlx5_lag *ldev, int idx, + u32 flags) +{ + struct lag_func *pf = mlx5_lag_pf(ldev, idx); + struct mlx5_eswitch *esw; + int ret; + + if (pf->dev->priv.flags & flags) + return 0; + + esw = pf->dev->priv.eswitch; + mlx5_esw_reps_block(esw); + ret = mlx5_eswitch_reload_ib_reps(esw); + mlx5_esw_reps_unblock(esw); + + return ret; +} + static int mlx5_lag_reload_ib_reps_unlocked(struct mlx5_lag *ldev, u32 flags, u32 filter, bool cont_on_fail) { - struct lag_func *pf; + int master_idx = mlx5_lag_get_dev_index_by_seq_filter(ldev, MLX5_LAG_P1, + filter); int ret; int i; + if (master_idx < 0) + return -EINVAL; + + ret = mlx5_lag_reload_ib_reps_idx(ldev, master_idx, flags); + if (ret && !cont_on_fail) + return ret; + mlx5_lag_for_each(i, 0, ldev, filter) { - pf = mlx5_lag_pf(ldev, i); - if (!(pf->dev->priv.flags & flags)) { - struct mlx5_eswitch *esw; - - esw = pf->dev->priv.eswitch; - mlx5_esw_reps_block(esw); - ret = mlx5_eswitch_reload_ib_reps(esw); - mlx5_esw_reps_unblock(esw); - if (ret && !cont_on_fail) - return ret; - } + if (i == master_idx) + continue; + ret = mlx5_lag_reload_ib_reps_idx(ldev, i, flags); + if (ret && !cont_on_fail) + return ret; } return 0; diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c b/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c index 6b4ad3c53f2f..424040918fa3 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/lag/shared_fdb.c @@ -270,6 +270,7 @@ int mlx5_lag_shared_fdb_create(struct mlx5_lag *ldev, pf->sd_fdb_active = false; } mlx5_lag_destroy_single_fdb_filter(ldev, group_id); + mlx5_lag_unload_reps_from_locked(ldev, filter); } err_add_devices: mlx5_lag_add_devices_filter(ldev, filter); diff --git a/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c b/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c index 64f92427602d..75855481522b 100644 --- a/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c +++ b/drivers/net/ethernet/mellanox/mlx5/core/lib/devcom.c @@ -37,6 +37,7 @@ struct mlx5_devcom_comp { struct mlx5_devcom_key key; mlx5_devcom_event_handler_t handler; struct kref ref; + int nr_devs; bool ready; struct rw_semaphore sem; struct lock_class_key lock_key; @@ -170,6 +171,7 @@ devcom_alloc_comp_dev(struct mlx5_devcom_dev *devc, down_write(&comp->sem); list_add_tail(&devcom->list, &comp->comp_dev_list_head); + WRITE_ONCE(comp->nr_devs, comp->nr_devs + 1); up_write(&comp->sem); return devcom; @@ -182,6 +184,7 @@ devcom_free_comp_dev(struct mlx5_devcom_comp_dev *devcom) down_write(&comp->sem); list_del(&devcom->list); + WRITE_ONCE(comp->nr_devs, comp->nr_devs - 1); up_write(&comp->sem); kref_put(&devcom->devc->ref, mlx5_devcom_dev_release); @@ -284,7 +287,7 @@ int mlx5_devcom_comp_get_size(struct mlx5_devcom_comp_dev *devcom) { struct mlx5_devcom_comp *comp = devcom->comp; - return kref_read(&comp->ref); + return READ_ONCE(comp->nr_devs); } int mlx5_devcom_locked_send_event(struct mlx5_devcom_comp_dev *devcom, diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_csr.h b/drivers/net/ethernet/meta/fbnic/fbnic_csr.h index 64b958df7774..baba3471bf5a 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_csr.h +++ b/drivers/net/ethernet/meta/fbnic/fbnic_csr.h @@ -974,6 +974,7 @@ enum { /* PUL User Registers */ #define FBNIC_CSR_START_PUL_USER 0x31000 /* CSR section delimiter */ #define FBNIC_PUL_OB_TLP_HDR_AW_CFG 0x3103d /* 0xc40f4 */ +#define FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH_MODE CSR_BIT(20) #define FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH CSR_BIT(19) #define FBNIC_PUL_OB_TLP_HDR_AW_CFG_BME CSR_BIT(18) #define FBNIC_PUL_OB_TLP_HDR_AW_CFG_RDE_ATTR CSR_GENMASK(17, 15) @@ -1215,6 +1216,10 @@ enum { #define FBNIC_IPC_MBX_DESC_LEN_MASK DESC_GENMASK(63, 48) #define FBNIC_IPC_MBX_DESC_EOM DESC_BIT(46) #define FBNIC_IPC_MBX_DESC_ADDR_MASK DESC_GENMASK(45, 3) +/* Set with FW_CMPL when the FW completed a descriptor without successfully + * processing it (e.g. a mailbox DMA error); the completion has no valid data. + */ +#define FBNIC_IPC_MBX_DESC_FW_ERR DESC_BIT(2) #define FBNIC_IPC_MBX_DESC_FW_CMPL DESC_BIT(1) #define FBNIC_IPC_MBX_DESC_HOST_CMPL DESC_BIT(0) diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c b/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c index 3c4563c8f403..6edfa0aa69f1 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_debugfs.c @@ -539,8 +539,8 @@ static void fbnic_dbg_fw_mbx_display(struct seq_file *s, /* Generate header */ seq_puts(s, mbx_idx == FBNIC_IPC_MBX_RX_IDX ? "Rx\n" : "Tx\n"); - seq_printf(s, "Rdy: %d Head: %d Tail: %d\n", - mbx->ready, mbx->head, mbx->tail); + seq_printf(s, "Rdy: %d Head: %d Tail: %d resp_error: %llu\n", + mbx->ready, mbx->head, mbx->tail, mbx->resp_error); snprintf(hdr, sizeof(hdr), "%3s %-4s %s %-12s %s %-3s %-16s\n", "Idx", "Len", "E", "Addr", "F", "H", "Raw"); diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c index 0e47088ec44b..76e9a545bb16 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_ethtool.c @@ -313,6 +313,11 @@ fbnic_get_ringparam(struct net_device *netdev, struct ethtool_ringparam *ring, kernel_ring->hds_thresh = fbn->hds_thresh; } +static u32 fbnic_ring_size_pow2(u32 size) +{ + return size ? roundup_pow_of_two(size) : 0; +} + static void fbnic_set_rings(struct fbnic_net *fbn, struct ethtool_ringparam *ring, struct kernel_ethtool_ringparam *kernel_ring) @@ -334,10 +339,10 @@ fbnic_set_ringparam(struct net_device *netdev, struct ethtool_ringparam *ring, struct fbnic_net *clone; int err; - ring->rx_pending = roundup_pow_of_two(ring->rx_pending); - ring->rx_mini_pending = roundup_pow_of_two(ring->rx_mini_pending); - ring->rx_jumbo_pending = roundup_pow_of_two(ring->rx_jumbo_pending); - ring->tx_pending = roundup_pow_of_two(ring->tx_pending); + ring->rx_pending = fbnic_ring_size_pow2(ring->rx_pending); + ring->rx_mini_pending = fbnic_ring_size_pow2(ring->rx_mini_pending); + ring->rx_jumbo_pending = fbnic_ring_size_pow2(ring->rx_jumbo_pending); + ring->tx_pending = fbnic_ring_size_pow2(ring->tx_pending); /* These are absolute minimums allowing the device and driver to operate * but not necessarily guarantee reasonable performance. Settings below @@ -2025,7 +2030,8 @@ static const struct ethtool_ops fbnic_ethtool_ops = { ETHTOOL_OP_NEEDS_RTNL_SPAUSEPARAM | ETHTOOL_OP_NEEDS_RTNL_SCHANNELS | ETHTOOL_OP_NEEDS_RTNL_SRINGPARAM | - ETHTOOL_OP_NEEDS_RTNL_GLINK, + ETHTOOL_OP_NEEDS_RTNL_GLINK | + ETHTOOL_OP_NEEDS_RTNL_TEST, .get_drvinfo = fbnic_get_drvinfo, .get_regs_len = fbnic_get_regs_len, .get_regs = fbnic_get_regs, diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_fw.c b/drivers/net/ethernet/meta/fbnic/fbnic_fw.c index 283d25fae79e..6d7eb8479edf 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_fw.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_fw.c @@ -60,8 +60,15 @@ static void fbnic_mbx_reset_desc_ring(struct fbnic_dev *fbd, int mbx_idx) */ switch (mbx_idx) { case FBNIC_IPC_MBX_RX_IDX: + /* Clearing BME blocks the device from writing to the host + * but leaves the requests parked in the write pipeline. The + * write path only clears outstanding requests when both FLUSH + * and FLUSH_MODE are set; FLUSH_MODE lets them drain without + * landing on the host. + */ wr32(fbd, FBNIC_PUL_OB_TLP_HDR_AW_CFG, - FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH); + FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH | + FBNIC_PUL_OB_TLP_HDR_AW_CFG_FLUSH_MODE); break; case FBNIC_IPC_MBX_TX_IDX: wr32(fbd, FBNIC_PUL_OB_TLP_HDR_AR_CFG, @@ -285,6 +292,12 @@ static void fbnic_mbx_process_tx_msgs(struct fbnic_dev *fbd) if (!(desc & FBNIC_IPC_MBX_DESC_FW_CMPL)) break; + if (desc & FBNIC_IPC_MBX_DESC_FW_ERR) { + tx_mbx->resp_error++; + dev_warn_ratelimited(fbd->dev, + "FW completed a Tx mailbox request with an error\n"); + } + fbnic_mbx_unmap_and_free_msg(fbd, FBNIC_IPC_MBX_TX_IDX, head); head++; @@ -1666,6 +1679,13 @@ static void fbnic_mbx_process_rx_msgs(struct fbnic_dev *fbd) if (!(desc & FBNIC_IPC_MBX_DESC_FW_CMPL)) break; + if (desc & FBNIC_IPC_MBX_DESC_FW_ERR) { + rx_mbx->resp_error++; + dev_warn_ratelimited(fbd->dev, + "FW reported an error on an Rx mailbox message; dropping\n"); + goto next_page; + } + dma_sync_single_for_cpu(fbd->dev, rx_mbx->buf_info[head].addr, FBNIC_RX_PAGE_SIZE, DMA_FROM_DEVICE); @@ -1733,7 +1753,9 @@ void fbnic_mbx_poll(struct fbnic_dev *fbd) int fbnic_mbx_poll_tx_ready(struct fbnic_dev *fbd) { struct fbnic_fw_mbx *tx_mbx = &fbd->mbx[FBNIC_IPC_MBX_TX_IDX]; + struct fbnic_fw_mbx *rx_mbx = &fbd->mbx[FBNIC_IPC_MBX_RX_IDX]; unsigned long timeout = jiffies + 10 * HZ + 1; + u64 tx_resp_error, rx_resp_error; int err, i; do { @@ -1764,6 +1786,9 @@ int fbnic_mbx_poll_tx_ready(struct fbnic_dev *fbd) * mgmt.version once we get the actual version from the firmware * in the capabilities request message. */ +send_cap_req: + tx_resp_error = tx_mbx->resp_error; + rx_resp_error = rx_mbx->resp_error; err = fbnic_fw_xmit_simple_msg(fbd, FBNIC_TLV_MSG_ID_HOST_CAP_REQ); if (err) goto clean_mbx; @@ -1781,9 +1806,27 @@ int fbnic_mbx_poll_tx_ready(struct fbnic_dev *fbd) msleep(20); fbnic_mbx_poll(fbd); + /* A valid capabilities response ends the poll. Check it + * before the FW_ERR retry below so a response parsed in the + * same poll as an unrelated FW_ERR is not discarded. + */ + if (fbd->fw_cap.running.mgmt.version >= MIN_FW_VER_CODE) + break; + /* set err, but wait till mgmt.version check to report it */ - if (!time_is_after_jiffies(timeout)) + if (!time_is_after_jiffies(timeout)) { err = -ETIMEDOUT; + continue; + } + + /* The FW can flag our capabilities request (Tx) or its + * response (Rx) with FW_ERR, in which case it produced no + * usable response. The ring is not wedged, so re-issue the + * request instead of spinning until the timeout. + */ + if (tx_mbx->resp_error != tx_resp_error || + rx_mbx->resp_error != rx_resp_error) + goto send_cap_req; } return 0; diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_fw.h b/drivers/net/ethernet/meta/fbnic/fbnic_fw.h index d84723e4cfa3..5f9969247e30 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_fw.h +++ b/drivers/net/ethernet/meta/fbnic/fbnic_fw.h @@ -13,6 +13,7 @@ struct fbnic_tlv_msg; struct fbnic_fw_mbx { u8 ready, head, tail; + u64 resp_error; struct { struct fbnic_tlv_msg *msg; dma_addr_t addr; diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_pci.c b/drivers/net/ethernet/meta/fbnic/fbnic_pci.c index 8b9bc9e8ea56..c6698e3002a1 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_pci.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_pci.c @@ -434,6 +434,7 @@ static int fbnic_pm_suspend(struct device *dev) { struct fbnic_dev *fbd = dev_get_drvdata(dev); struct net_device *netdev = fbd->netdev; + struct fbnic_net *fbn; if (fbnic_init_failure(fbd)) goto null_uc_addr; @@ -441,11 +442,16 @@ static int fbnic_pm_suspend(struct device *dev) rtnl_lock(); netdev_lock(netdev); + fbn = netdev_priv(netdev); + netif_device_detach(netdev); if (netif_running(netdev)) netdev->netdev_ops->ndo_stop(netdev); + /* The IRQs are about to be freed, so drop the napi vector count */ + fbn->num_napi = 0; + netdev_unlock(netdev); rtnl_unlock(); @@ -508,16 +514,20 @@ static int __fbnic_pm_resume(struct device *dev) if (fbnic_init_failure(fbd)) return 0; + rtnl_lock(); + netdev_lock(netdev); + fbn = netdev_priv(netdev); /* Reset the queues if needed */ fbnic_reset_queues(fbn, fbn->num_tx_queues, fbn->num_rx_queues); - rtnl_lock(); - netdev_lock(netdev); - - if (netif_running(netdev)) + if (netif_running(netdev)) { err = __fbnic_open(fbn); + /* On failure the vectors are freed, so drop the count */ + if (err) + fbn->num_napi = 0; + } netdev_unlock(netdev); rtnl_unlock(); diff --git a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c index e7918d3f6aba..10caacffee0f 100644 --- a/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c +++ b/drivers/net/ethernet/meta/fbnic/fbnic_txrx.c @@ -7,6 +7,7 @@ #include #include #include +#include #include #include #include @@ -1622,7 +1623,7 @@ fbnic_alloc_qt_page_pools(struct fbnic_net *fbn, struct fbnic_q_triad *qt, return 0; err_destroy_sub0: - page_pool_destroy(pp); + page_pool_destroy(qt->sub0.page_pool); return PTR_ERR(pp); } @@ -1791,7 +1792,7 @@ int fbnic_alloc_napi_vectors(struct fbnic_net *fbn) int err; /* Allocate 1 Tx queue per napi vector */ - if (num_napi < FBNIC_MAX_TXQS && num_napi == num_tx + num_rx) { + if (num_napi <= FBNIC_MAX_TXQS && num_napi == num_tx + num_rx) { while (num_tx) { err = fbnic_alloc_napi_vector(fbd, fbn, num_napi, v_idx, @@ -2853,6 +2854,17 @@ void fbnic_napi_depletion_check(struct net_device *netdev) fbnic_wrfl(fbd); } +/* Returns the napi vector servicing an Rx queue, or NULL if the datapath + * is torn down. The association is published by fbnic_set_netif_napi() + * and cleared by fbnic_reset_netif_napi(), both under the instance lock. + */ +static struct fbnic_napi_vector *fbnic_rxq_nv(struct net_device *dev, int idx) +{ + struct napi_struct *napi = __netif_get_rx_queue(dev, idx)->napi; + + return napi ? container_of(napi, struct fbnic_napi_vector, napi) : NULL; +} + static int fbnic_queue_mem_alloc(struct net_device *dev, struct netdev_queue_config *qcfg, void *qmem, int idx) @@ -2865,8 +2877,16 @@ static int fbnic_queue_mem_alloc(struct net_device *dev, if (!netif_running(dev)) return fbnic_alloc_qt_page_pools(fbn, qt, idx); + /* A failed PCIe recovery or resume can leave the datapath torn down + * while netif_running() is still true. This ndo runs before + * netdev_rx_queue_restart() checks netif_running(), so bail out + * rather than touching rings and vectors that are already freed. + */ + nv = fbnic_rxq_nv(dev, idx); + if (!nv) + return -ENETDOWN; + real = container_of(fbn->rx[idx], struct fbnic_q_triad, cmpl); - nv = fbn->napi[idx % fbn->num_napi]; fbnic_ring_init(&qt->sub0, real->sub0.doorbell, real->sub0.q_idx, real->sub0.flags); @@ -2917,7 +2937,7 @@ static int fbnic_queue_start(struct net_device *dev, struct fbnic_q_triad *real; real = container_of(fbn->rx[idx], struct fbnic_q_triad, cmpl); - nv = fbn->napi[idx % fbn->num_napi]; + nv = fbnic_rxq_nv(dev, idx); fbnic_aggregate_ring_bdq_counters(fbn, &real->sub0); fbnic_aggregate_ring_bdq_counters(fbn, &real->sub1); @@ -2939,7 +2959,7 @@ static int fbnic_queue_stop(struct net_device *dev, void *qmem, int idx) int err; real = container_of(fbn->rx[idx], struct fbnic_q_triad, cmpl); - nv = fbn->napi[idx % fbn->num_napi]; + nv = fbnic_rxq_nv(dev, idx); fbnic_dbg_nv_exit(nv); napi_disable_locked(&nv->napi); diff --git a/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c b/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c index 9e7c285eaa6b..960d7513aa8d 100644 --- a/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c +++ b/drivers/net/ethernet/netronome/nfp/crypto/ipsec.c @@ -625,11 +625,12 @@ int nfp_net_ipsec_rx(struct nfp_meta_parsed *meta, struct sk_buff *skb) xa_lock(&nn->xa_ipsec); x = xa_load(&nn->xa_ipsec, saidx); + if (x) + xfrm_state_hold(x); xa_unlock(&nn->xa_ipsec); if (!x) return -EINVAL; - xfrm_state_hold(x); sp->xvec[sp->len++] = x; sp->olen++; xo = xfrm_offload(skb); diff --git a/drivers/net/ethernet/spacemit/k1_emac.c b/drivers/net/ethernet/spacemit/k1_emac.c index f7f16397a2c2..d641ac26a1e8 100644 --- a/drivers/net/ethernet/spacemit/k1_emac.c +++ b/drivers/net/ethernet/spacemit/k1_emac.c @@ -803,6 +803,9 @@ static void emac_tx_mem_map(struct emac_priv *priv, struct sk_buff *skb) while (i != head) { emac_free_tx_buf(priv, i); + tx_desc_addr = &((struct emac_desc *)tx_ring->desc_addr)[i]; + memset(tx_desc_addr, 0, sizeof(*tx_desc_addr)); + if (++i == tx_ring->total_cnt) i = 0; } diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c b/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c index 8d7042e68926..72bdbcb5e863 100644 --- a/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c +++ b/drivers/net/ethernet/stmicro/stmmac/dwmac-rk.c @@ -1162,8 +1162,11 @@ static int gmac_clk_enable(struct rk_priv_data *bsp_priv, bool enable) return ret; ret = clk_prepare_enable(bsp_priv->clk_phy); - if (ret) + if (ret) { + clk_bulk_disable_unprepare(bsp_priv->num_clks, + bsp_priv->clks); return ret; + } rk_configure_io_clksel(bsp_priv); rk_ungate_rmii_clock(bsp_priv); diff --git a/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c b/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c index 2994df41ec2c..c6a8f8d73501 100644 --- a/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c +++ b/drivers/net/ethernet/stmicro/stmmac/dwmac4_descs.c @@ -474,11 +474,11 @@ static void dwmac4_set_sarc(struct dma_desc *p, u32 sarc_type) sarc_type)); } -static int set_16kib_bfsize(int mtu) +static int set_16kib_bfsize(int len) { int ret = 0; - if (unlikely(mtu >= BUF_SIZE_8KiB)) + if (unlikely(len > BUF_SIZE_8KiB)) ret = BUF_SIZE_16KiB; return ret; } diff --git a/drivers/net/ethernet/stmicro/stmmac/hwif.h b/drivers/net/ethernet/stmicro/stmmac/hwif.h index 9314bcb85c22..857f7562c6c6 100644 --- a/drivers/net/ethernet/stmicro/stmmac/hwif.h +++ b/drivers/net/ethernet/stmicro/stmmac/hwif.h @@ -540,7 +540,7 @@ struct stmmac_mode_ops { bool (*is_jumbo_frm)(unsigned int len, bool enh_desc); int (*jumbo_frm)(struct stmmac_tx_queue *tx_q, struct sk_buff *skb, int csum); - int (*set_16kib_bfsize)(int mtu); + int (*set_16kib_bfsize)(int len); void (*init_desc3)(struct dma_desc *p); void (*refill_desc3)(struct stmmac_rx_queue *rx_q, struct dma_desc *p); void (*clean_desc3)(struct stmmac_tx_queue *tx_q, struct dma_desc *p); diff --git a/drivers/net/ethernet/stmicro/stmmac/ring_mode.c b/drivers/net/ethernet/stmicro/stmmac/ring_mode.c index 78fc6aa5bbe9..dd796cb74419 100644 --- a/drivers/net/ethernet/stmicro/stmmac/ring_mode.c +++ b/drivers/net/ethernet/stmicro/stmmac/ring_mode.c @@ -123,10 +123,10 @@ static void clean_desc3(struct stmmac_tx_queue *tx_q, struct dma_desc *p) p->des3 = 0; } -static int set_16kib_bfsize(int mtu) +static int set_16kib_bfsize(int len) { int ret = 0; - if (unlikely(mtu > BUF_SIZE_8KiB)) + if (unlikely(len > BUF_SIZE_8KiB)) ret = BUF_SIZE_16KiB; return ret; } diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c index 5f717f02c3c0..29025eb41698 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_main.c @@ -1536,17 +1536,17 @@ static unsigned int stmmac_rx_offset(struct stmmac_priv *priv) return NET_SKB_PAD + NET_IP_ALIGN; } -static int stmmac_set_bfsize(int mtu) +static int stmmac_set_bfsize(int len) { int ret; - if (mtu >= BUF_SIZE_8KiB) + if (len > BUF_SIZE_8KiB) ret = BUF_SIZE_16KiB; - else if (mtu >= BUF_SIZE_4KiB) + else if (len > BUF_SIZE_4KiB) ret = BUF_SIZE_8KiB; - else if (mtu >= BUF_SIZE_2KiB) + else if (len > BUF_SIZE_2KiB) ret = BUF_SIZE_4KiB; - else if (mtu > DEFAULT_BUFSIZE) + else if (len > DEFAULT_BUFSIZE) ret = BUF_SIZE_2KiB; else ret = DEFAULT_BUFSIZE; @@ -4063,7 +4063,7 @@ static struct stmmac_dma_conf * stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu) { struct stmmac_dma_conf *dma_conf; - int bfsize, ret; + int bfsize, len, ret; u8 chan; dma_conf = kzalloc_obj(*dma_conf); @@ -4073,13 +4073,15 @@ stmmac_setup_dma_desc(struct stmmac_priv *priv, unsigned int mtu) return ERR_PTR(-ENOMEM); } - /* Returns 0 or BUF_SIZE_16KiB if mtu > 8KiB and dwmac4 or ring mode */ - bfsize = stmmac_set_16kib_bfsize(priv, mtu); + len = mtu + ETH_HLEN + 2 * VLAN_HLEN + ETH_FCS_LEN; + + /* Returns 0 or BUF_SIZE_16KiB if len > 8KiB and dwmac4 or ring mode */ + bfsize = stmmac_set_16kib_bfsize(priv, len); if (bfsize < 0) bfsize = 0; if (bfsize < BUF_SIZE_16KiB) - bfsize = stmmac_set_bfsize(mtu); + bfsize = stmmac_set_bfsize(len); dma_conf->dma_buf_sz = bfsize; /* Chose the tx/rx size from the already defined one in the @@ -5883,6 +5885,7 @@ static int stmmac_rx(struct stmmac_priv *priv, int limit, u32 queue) if (!skb) { page_pool_recycle_direct(rx_q->page_pool, buf->page); + buf->page = NULL; rx_dropped++; count++; goto drain_data; diff --git a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c index f97f32369e90..75647992cbb4 100644 --- a/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c +++ b/drivers/net/ethernet/stmicro/stmmac/stmmac_selftests.c @@ -13,6 +13,7 @@ #include #include #include +#include #include #include #include @@ -30,6 +31,7 @@ struct stmmachdr { sizeof(struct stmmachdr)) #define STMMAC_TEST_PKT_MAGIC 0xdeadcafecafedeadULL #define STMMAC_LB_TIMEOUT msecs_to_jiffies(200) +#define STMMAC_SFT_MAX_LPI (5 * USEC_PER_SEC) struct stmmac_packet_attrs { int vlan; @@ -238,6 +240,10 @@ struct stmmac_test_priv { struct stmmac_packet_attrs *packet; struct packet_type pt; struct completion comp; + __be16 packet_type; + int (*func)(struct sk_buff *skb, struct net_device *ndev, + struct packet_type *pt, struct net_device *orig_ndev); + bool capture_all; int double_vlan; int vlan_id; int ok; @@ -317,6 +323,52 @@ static int stmmac_test_loopback_validate(struct sk_buff *skb, return 0; } +static int stmmac_sft_filter(struct sk_buff *skb, struct net_device *ndev, + struct packet_type *pt, + struct net_device *orig_ndev) +{ + struct stmmac_test_priv *tpriv = pt->af_packet_priv; + struct ethhdr *hdr = eth_hdr(skb); + int ret = 0; + + if (hdr->h_proto == tpriv->packet_type) { + struct sk_buff *nskb = skb_clone(skb, GFP_ATOMIC); + + if (nskb) + ret = tpriv->func(nskb, ndev, pt, orig_ndev); + } + + kfree_skb(skb); + return ret; +} + +static void stmmac_sft_add_pack(struct packet_type *pt) +{ + struct stmmac_test_priv *tpriv = pt->af_packet_priv; + + if (netdev_uses_dsa(tpriv->pt.dev) || tpriv->capture_all) { + tpriv->packet_type = tpriv->pt.type; + tpriv->func = tpriv->pt.func; + + /* DSA conduit will report ETH_P_XDSA, so our packet handler + * won't match. Let's register a ETH_P_ALL match and filter + * manually in stmmac_sft_filter. This is also useful for + * VLAN tests, to capture packets otherwise marked as + * OTHERHOST. + */ + tpriv->pt.type = htons(ETH_P_ALL); + tpriv->pt.func = stmmac_sft_filter; + tpriv->pt.ignore_outgoing = true; + } + + dev_add_pack(pt); +} + +static void stmmac_sft_remove_pack(struct packet_type *pt) +{ + dev_remove_pack(pt); +} + static int __stmmac_test_loopback(struct stmmac_priv *priv, struct stmmac_packet_attrs *attr) { @@ -338,7 +390,7 @@ static int __stmmac_test_loopback(struct stmmac_priv *priv, tpriv->packet = attr; if (!attr->dont_wait) - dev_add_pack(&tpriv->pt); + stmmac_sft_add_pack(&tpriv->pt); skb = stmmac_test_get_udp_skb(priv, attr); if (!skb) { @@ -361,7 +413,7 @@ static int __stmmac_test_loopback(struct stmmac_priv *priv, cleanup: if (!attr->dont_wait) - dev_remove_pack(&tpriv->pt); + stmmac_sft_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -434,12 +486,16 @@ static int stmmac_test_mmc(struct stmmac_priv *priv) static int stmmac_test_eee(struct stmmac_priv *priv) { struct stmmac_extra_stats *initial, *final; - int retries = 10; + unsigned long timeout, max_duration; int ret; if (!priv->dma_cap.eee || !priv->eee_active) return -EOPNOTSUPP; + /* Bail out if the configured LPI timer is too long */ + if (priv->tx_lpi_timer > STMMAC_SFT_MAX_LPI) + return -EOPNOTSUPP; + initial = kzalloc_obj(*initial); if (!initial) return -ENOMEM; @@ -450,14 +506,21 @@ static int stmmac_test_eee(struct stmmac_priv *priv) goto out_free_initial; } + /* Snapshot stats, we want to count the in_lpi events. We may enter + * LPI just after the packet was sent. + */ memcpy(initial, &priv->xstats, sizeof(*initial)); + /* Send a frame, then wait to enter LPI */ ret = stmmac_test_mac_loopback(priv); if (ret) goto out_free_final; + max_duration = usecs_to_jiffies(2 * priv->tx_lpi_timer); + /* We have no traffic in the line so, sooner or later it will go LPI */ - while (--retries) { + timeout = jiffies + max_duration; + while (!time_after(jiffies, timeout)) { memcpy(final, &priv->xstats, sizeof(*final)); if (final->irq_tx_path_in_lpi_mode_n > @@ -466,20 +529,38 @@ static int stmmac_test_eee(struct stmmac_priv *priv) msleep(100); } - if (!retries) { + memcpy(final, &priv->xstats, sizeof(*final)); + if (final->irq_tx_path_in_lpi_mode_n <= + initial->irq_tx_path_in_lpi_mode_n) { ret = -ETIMEDOUT; goto out_free_final; } - if (final->irq_tx_path_in_lpi_mode_n <= - initial->irq_tx_path_in_lpi_mode_n) { - ret = -EINVAL; + /* Re-snapshot, as we want to measure exit_lpi events. We should be + * in LPI right now. + */ + memcpy(initial, &priv->xstats, sizeof(*initial)); + + /* TX something so we go out of LPI */ + ret = stmmac_test_mac_loopback(priv); + if (ret) goto out_free_final; + + /* Wait for the exit LPI interrupt */ + timeout = jiffies + max_duration; + while (!time_after(jiffies, timeout)) { + memcpy(final, &priv->xstats, sizeof(*final)); + + if (final->irq_tx_path_exit_lpi_mode_n > + initial->irq_tx_path_exit_lpi_mode_n) + break; + msleep(100); } + memcpy(final, &priv->xstats, sizeof(*final)); if (final->irq_tx_path_exit_lpi_mode_n <= initial->irq_tx_path_exit_lpi_mode_n) { - ret = -EINVAL; + ret = -ETIMEDOUT; goto out_free_final; } @@ -787,7 +868,7 @@ static int stmmac_test_flowctrl(struct stmmac_priv *priv) tpriv->pt.func = stmmac_test_flowctrl_validate; tpriv->pt.dev = priv->dev; tpriv->pt.af_packet_priv = tpriv; - dev_add_pack(&tpriv->pt); + stmmac_sft_add_pack(&tpriv->pt); /* Compute minimum number of packets to make FIFO full */ pkt_count = rx_fifo_size; @@ -843,7 +924,7 @@ static int stmmac_test_flowctrl(struct stmmac_priv *priv) cleanup: dev_mc_del(priv->dev, paddr); dev_set_promiscuity(priv->dev, -1); - dev_remove_pack(&tpriv->pt); + stmmac_sft_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -885,6 +966,11 @@ static int stmmac_test_vlan_validate(struct sk_buff *skb, goto out; if (skb_headlen(skb) < (STMMAC_TEST_PKT_SIZE - ETH_HLEN)) goto out; + + ehdr = (struct ethhdr *)skb_mac_header(skb); + if (!ether_addr_equal_unaligned(ehdr->h_dest, tpriv->packet->dst)) + goto out; + if (tpriv->vlan_id) { if (skb->vlan_proto != htons(proto)) goto out; @@ -896,10 +982,6 @@ static int stmmac_test_vlan_validate(struct sk_buff *skb, } } - ehdr = (struct ethhdr *)skb_mac_header(skb); - if (!ether_addr_equal_unaligned(ehdr->h_dest, tpriv->packet->dst)) - goto out; - ihdr = ip_hdr(skb); if (tpriv->double_vlan) ihdr = (struct iphdr *)(skb_network_header(skb) + 4); @@ -941,6 +1023,7 @@ static int __stmmac_test_vlanfilt(struct stmmac_priv *priv) tpriv->pt.dev = priv->dev; tpriv->pt.af_packet_priv = tpriv; tpriv->packet = &attr; + tpriv->capture_all = true; /* * As we use HASH filtering, false positives may appear. This is a @@ -948,18 +1031,20 @@ static int __stmmac_test_vlanfilt(struct stmmac_priv *priv) * HASH values. */ tpriv->vlan_id = 0x123; - dev_add_pack(&tpriv->pt); ret = vlan_vid_add(priv->dev, htons(ETH_P_8021Q), tpriv->vlan_id); if (ret) goto cleanup; + attr.vlan = 1; + attr.dst = priv->dev->dev_addr; + attr.sport = 9; + attr.dport = 9; + + stmmac_sft_add_pack(&tpriv->pt); + for (i = 0; i < 4; i++) { - attr.vlan = 1; attr.vlan_id_out = tpriv->vlan_id + i; - attr.dst = priv->dev->dev_addr; - attr.sport = 9; - attr.dport = 9; skb = stmmac_test_get_udp_skb(priv, &attr); if (!skb) { @@ -986,9 +1071,9 @@ static int __stmmac_test_vlanfilt(struct stmmac_priv *priv) } vlan_del: + stmmac_sft_remove_pack(&tpriv->pt); vlan_vid_del(priv->dev, htons(ETH_P_8021Q), tpriv->vlan_id); cleanup: - dev_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -1035,6 +1120,7 @@ static int __stmmac_test_dvlanfilt(struct stmmac_priv *priv) tpriv->pt.dev = priv->dev; tpriv->pt.af_packet_priv = tpriv; tpriv->packet = &attr; + tpriv->capture_all = true; /* * As we use HASH filtering, false positives may appear. This is a @@ -1042,18 +1128,20 @@ static int __stmmac_test_dvlanfilt(struct stmmac_priv *priv) * HASH values. */ tpriv->vlan_id = 0x123; - dev_add_pack(&tpriv->pt); ret = vlan_vid_add(priv->dev, htons(ETH_P_8021AD), tpriv->vlan_id); if (ret) goto cleanup; + attr.vlan = 2; + attr.dst = priv->dev->dev_addr; + attr.sport = 9; + attr.dport = 9; + + stmmac_sft_add_pack(&tpriv->pt); + for (i = 0; i < 4; i++) { - attr.vlan = 2; attr.vlan_id_out = tpriv->vlan_id + i; - attr.dst = priv->dev->dev_addr; - attr.sport = 9; - attr.dport = 9; skb = stmmac_test_get_udp_skb(priv, &attr); if (!skb) { @@ -1080,9 +1168,9 @@ static int __stmmac_test_dvlanfilt(struct stmmac_priv *priv) } vlan_del: + stmmac_sft_remove_pack(&tpriv->pt); vlan_vid_del(priv->dev, htons(ETH_P_8021AD), tpriv->vlan_id); cleanup: - dev_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -1313,7 +1401,7 @@ static int stmmac_test_vlanoff_common(struct stmmac_priv *priv, bool svlan) tpriv->pt.af_packet_priv = tpriv; tpriv->packet = &attr; tpriv->vlan_id = 0x123; - dev_add_pack(&tpriv->pt); + tpriv->capture_all = true; ret = vlan_vid_add(priv->dev, htons(proto), tpriv->vlan_id); if (ret) @@ -1321,6 +1409,8 @@ static int stmmac_test_vlanoff_common(struct stmmac_priv *priv, bool svlan) attr.dst = priv->dev->dev_addr; + stmmac_sft_add_pack(&tpriv->pt); + skb = stmmac_test_get_udp_skb(priv, &attr); if (!skb) { ret = -ENOMEM; @@ -1338,9 +1428,9 @@ static int stmmac_test_vlanoff_common(struct stmmac_priv *priv, bool svlan) ret = tpriv->ok ? 0 : -ETIMEDOUT; vlan_del: + stmmac_sft_remove_pack(&tpriv->pt); vlan_vid_del(priv->dev, htons(proto), tpriv->vlan_id); cleanup: - dev_remove_pack(&tpriv->pt); kfree(tpriv); return ret; } @@ -1352,7 +1442,7 @@ static int stmmac_test_vlanoff(struct stmmac_priv *priv) static int stmmac_test_svlanoff(struct stmmac_priv *priv) { - if (!priv->dma_cap.dvlan) + if (!(priv->dev->features & NETIF_F_HW_VLAN_STAG_TX)) return -EOPNOTSUPP; return stmmac_test_vlanoff_common(priv, true); } @@ -1719,6 +1809,9 @@ static int __stmmac_test_jumbo(struct stmmac_priv *priv, u16 queue) struct stmmac_packet_attrs attr = { }; int size = priv->dma_conf.dma_buf_sz; + if (!dwmac_is_xmac(priv->plat->core_type)) + size -= NET_IP_ALIGN; + attr.dst = priv->dev->dev_addr; attr.max_size = size - ETH_FCS_LEN; attr.queue_mapping = queue; diff --git a/drivers/net/ethernet/ti/netcp_core.c b/drivers/net/ethernet/ti/netcp_core.c index eb8fc2ed05f4..4f9e20468bbb 100644 --- a/drivers/net/ethernet/ti/netcp_core.c +++ b/drivers/net/ethernet/ti/netcp_core.c @@ -2225,7 +2225,7 @@ static int netcp_probe(struct platform_device *pdev) return -ENOMEM; pm_runtime_enable(&pdev->dev); - ret = pm_runtime_get_sync(&pdev->dev); + ret = pm_runtime_resume_and_get(&pdev->dev); if (ret < 0) { dev_err(dev, "Failed to enable NETCP power-domain\n"); pm_runtime_disable(&pdev->dev); diff --git a/drivers/net/ethernet/wangxun/libwx/wx_hw.c b/drivers/net/ethernet/wangxun/libwx/wx_hw.c index 260e14d5d541..9119c931b3a3 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_hw.c +++ b/drivers/net/ethernet/wangxun/libwx/wx_hw.c @@ -2517,6 +2517,7 @@ int wx_sw_init(struct wx *wx) } spin_lock_init(&wx->hw_stats_lock); + spin_lock_init(&wx->ptp_tx_lock); mutex_init(&wx->reset_lock); bitmap_zero(wx->state, WX_STATE_NBITS); bitmap_zero(wx->flags, WX_PF_FLAGS_NBITS); diff --git a/drivers/net/ethernet/wangxun/libwx/wx_lib.c b/drivers/net/ethernet/wangxun/libwx/wx_lib.c index 5d99e870de5e..287172348b07 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_lib.c +++ b/drivers/net/ethernet/wangxun/libwx/wx_lib.c @@ -1163,9 +1163,11 @@ static int wx_tx_map(struct wx_ring *tx_ring, i--; } - dev_kfree_skb_any(first->skb); - first->skb = NULL; - + /* first->skb is released by the caller, which keeps a reference on it + * until the PTP cleanup has compared it against wx->ptp_tx_skb. That + * prevents the address from being reused by a newer request while the + * comparison is pending. + */ tx_ring->next_to_use = i; return -ENOMEM; @@ -1612,9 +1614,11 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb, if (unlikely(skb_shinfo(skb)->tx_flags & SKBTX_HW_TSTAMP) && wx->ptp_clock) { + unsigned long flags; + + spin_lock_irqsave(&wx->ptp_tx_lock, flags); if (wx->tstamp_config.tx_type == HWTSTAMP_TX_ON && - !test_and_set_bit_lock(WX_STATE_PTP_TX_IN_PROGRESS, - wx->state)) { + !test_and_set_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state)) { skb_shinfo(skb)->tx_flags |= SKBTX_IN_PROGRESS; tx_flags |= WX_TX_FLAGS_TSTAMP; wx->ptp_tx_skb = skb_get(skb); @@ -1622,6 +1626,7 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb, } else { wx->tx_hwtstamp_skipped++; } + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); } /* record initial flags and protocol */ @@ -1640,19 +1645,35 @@ static netdev_tx_t wx_xmit_frame_ring(struct sk_buff *skb, wx->atr(tx_ring, first, ptype); if (wx_tx_map(tx_ring, first, hdr_len)) - goto cleanup_tx_tstamp; + goto out_drop; return NETDEV_TX_OK; out_drop: - dev_kfree_skb_any(first->skb); - first->skb = NULL; -cleanup_tx_tstamp: + /* The frame never reached the hardware, so no timestamp will ever be + * reported for it and the request has to be cancelled. The slot is + * shared, though: wx_ptp_clear_tx_timestamp() or wx_ptp_tx_hang() may + * have dropped our request already, and a transmit on another queue + * can have claimed the slot since. Only cancel it while it is still + * ours, otherwise we would free somebody else's skb and release their + * in-progress bit. + */ if (unlikely(tx_flags & WX_TX_FLAGS_TSTAMP)) { - dev_kfree_skb_any(wx->ptp_tx_skb); - wx->ptp_tx_skb = NULL; - wx->tx_hwtstamp_errors++; - clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + struct sk_buff *ptp_tx_skb = NULL; + unsigned long flags; + + spin_lock_irqsave(&wx->ptp_tx_lock, flags); + if (wx->ptp_tx_skb == skb) { + ptp_tx_skb = wx->ptp_tx_skb; + wx->ptp_tx_skb = NULL; + clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + wx->tx_hwtstamp_errors++; + } + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + dev_kfree_skb_any(ptp_tx_skb); } + dev_kfree_skb_any(first->skb); + first->skb = NULL; return NETDEV_TX_OK; } @@ -3348,5 +3369,23 @@ void wx_service_timer(struct timer_list *t) } EXPORT_SYMBOL(wx_service_timer); +void wx_soft_quiesce(struct wx *wx) +{ + if (!netif_running(wx->netdev) || + test_and_set_bit(WX_STATE_DOWN, wx->state)) + return; + + pci_clear_master(wx->pdev); + netif_tx_stop_all_queues(wx->netdev); + netif_carrier_off(wx->netdev); + netif_tx_disable(wx->netdev); + wx_napi_disable_all(wx); + wx_ptp_quiesce(wx); + + clear_bit(WX_FLAG_NEED_DO_RESET, wx->flags); + timer_delete_sync(&wx->service_timer); +} +EXPORT_SYMBOL(wx_soft_quiesce); + MODULE_DESCRIPTION("Common library for Wangxun(R) Ethernet drivers."); MODULE_LICENSE("GPL"); diff --git a/drivers/net/ethernet/wangxun/libwx/wx_lib.h b/drivers/net/ethernet/wangxun/libwx/wx_lib.h index aed6ea8cf0d6..11bd79985e17 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_lib.h +++ b/drivers/net/ethernet/wangxun/libwx/wx_lib.h @@ -41,5 +41,6 @@ void wx_set_ring(struct wx *wx, u32 new_tx_count, void wx_service_event_schedule(struct wx *wx); void wx_service_event_complete(struct wx *wx); void wx_service_timer(struct timer_list *t); +void wx_soft_quiesce(struct wx *wx); #endif /* _WX_LIB_H_ */ diff --git a/drivers/net/ethernet/wangxun/libwx/wx_ptp.c b/drivers/net/ethernet/wangxun/libwx/wx_ptp.c index 1165518d5522..65b8937f6e94 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_ptp.c +++ b/drivers/net/ethernet/wangxun/libwx/wx_ptp.c @@ -129,6 +129,34 @@ static int wx_ptp_settime64(struct ptp_clock_info *ptp, return 0; } +/** + * __wx_ptp_detach_tx_skb - detach the skb tracking the Tx timestamp request + * @wx: the private board structure + * + * Detach the skb of the outstanding request and release the in-progress bit, + * so that a new request can be submitted. + * + * This performs no register access. Callers that need a timestamp the hardware + * may have left latched must unlatch it themselves, while the device is known + * to be alive. wx_ptp_quiesce() runs during PCIe error recovery, where MMIO is + * not reliable, and therefore deliberately skips the unlatch. + * + * Context: Expects wx->ptp_tx_lock to be held by the caller. + * Return: the detached skb, or NULL if no request was outstanding. The caller + * owns the returned reference and must release it once the lock is dropped. + */ +static struct sk_buff *__wx_ptp_detach_tx_skb(struct wx *wx) +{ + struct sk_buff *skb = wx->ptp_tx_skb; + + lockdep_assert_held(&wx->ptp_tx_lock); + + wx->ptp_tx_skb = NULL; + clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + + return skb; +} + /** * wx_ptp_clear_tx_timestamp - utility function to clear Tx timestamp state * @wx: the private board structure @@ -139,12 +167,16 @@ static int wx_ptp_settime64(struct ptp_clock_info *ptp, */ static void wx_ptp_clear_tx_timestamp(struct wx *wx) { + struct sk_buff *skb; + unsigned long flags; + + spin_lock_irqsave(&wx->ptp_tx_lock, flags); + /* Unlatch a timestamp the hardware may have left pending. */ rd32ptp(wx, WX_TSC_1588_STMPH); - if (wx->ptp_tx_skb) { - dev_kfree_skb_any(wx->ptp_tx_skb); - wx->ptp_tx_skb = NULL; - } - clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + skb = __wx_ptp_detach_tx_skb(wx); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + dev_kfree_skb_any(skb); } /** @@ -175,49 +207,54 @@ static void wx_ptp_convert_to_hwtstamp(struct wx *wx, } /** - * wx_ptp_tx_hwtstamp - utility function which checks for TX time stamp + * wx_ptp_tx_hwtstamp_work - check for a pending Tx time stamp * @wx: the private board struct * - * if the timestamp is valid, we convert it into the timecounter ns - * value, then store that result into the shhwtstamps structure which - * is passed up the network stack + * If a Tx timestamp request is outstanding and the hardware has latched a + * valid value, we convert it into the timecounter ns value, then store that + * result into the shhwtstamps structure which is passed up the network stack. + * + * Return: 0 when there is nothing left to poll for, -1 when the timestamp is + * not available yet and the caller should poll again. */ -static void wx_ptp_tx_hwtstamp(struct wx *wx) +static int wx_ptp_tx_hwtstamp_work(struct wx *wx) { struct skb_shared_hwtstamps shhwtstamps; - struct sk_buff *skb = wx->ptp_tx_skb; + unsigned long flags; + struct sk_buff *skb; + u32 tsynctxctl; u64 regval = 0; - regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPL); - regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPH) << 32; - - wx_ptp_convert_to_hwtstamp(wx, &shhwtstamps, regval); - - wx->ptp_tx_skb = NULL; - clear_bit_unlock(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); - skb_tstamp_tx(skb, &shhwtstamps); - dev_kfree_skb_any(skb); - wx->tx_hwtstamp_pkts++; -} - -static int wx_ptp_tx_hwtstamp_work(struct wx *wx) -{ - u32 tsynctxctl; + spin_lock_irqsave(&wx->ptp_tx_lock, flags); /* we have to have a valid skb to poll for a timestamp */ if (!wx->ptp_tx_skb) { - wx_ptp_clear_tx_timestamp(wx); + rd32ptp(wx, WX_TSC_1588_STMPH); + __wx_ptp_detach_tx_skb(wx); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); return 0; } /* stop polling once we have a valid timestamp */ tsynctxctl = rd32ptp(wx, WX_TSC_1588_CTL); - if (tsynctxctl & WX_TSC_1588_CTL_VALID) { - wx_ptp_tx_hwtstamp(wx); - return 0; + if (!(tsynctxctl & WX_TSC_1588_CTL_VALID)) { + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + return -1; } - return -1; + regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPL); + regval |= (u64)rd32ptp(wx, WX_TSC_1588_STMPH) << 32; + skb = wx->ptp_tx_skb; + wx->ptp_tx_skb = NULL; + clear_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + wx_ptp_convert_to_hwtstamp(wx, &shhwtstamps, regval); + skb_tstamp_tx(skb, &shhwtstamps); + dev_kfree_skb_any(skb); + wx->tx_hwtstamp_pkts++; + + return 0; } /** @@ -296,24 +333,29 @@ static void wx_ptp_rx_hang(struct wx *wx) */ static void wx_ptp_tx_hang(struct wx *wx) { - bool timeout = time_is_before_jiffies(wx->ptp_tx_start + - WX_PTP_TX_TIMEOUT); - - if (!wx->ptp_tx_skb) - return; + struct sk_buff *skb = NULL; + unsigned long flags; - if (!test_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state)) - return; + spin_lock_irqsave(&wx->ptp_tx_lock, flags); /* If we haven't received a timestamp within the timeout, it is * reasonable to assume that it will never occur, so we can unlock the * timestamp bit when this occurs. */ - if (timeout) { - wx_ptp_clear_tx_timestamp(wx); - wx->tx_hwtstamp_timeouts++; - dev_warn(&wx->pdev->dev, "clearing Tx timestamp hang\n"); + if (wx->ptp_tx_skb && + test_bit(WX_STATE_PTP_TX_IN_PROGRESS, wx->state) && + time_is_before_jiffies(wx->ptp_tx_start + WX_PTP_TX_TIMEOUT)) { + rd32ptp(wx, WX_TSC_1588_STMPH); + skb = __wx_ptp_detach_tx_skb(wx); } + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + + if (!skb) + return; + + dev_kfree_skb_any(skb); + wx->tx_hwtstamp_timeouts++; + dev_warn(&wx->pdev->dev, "clearing Tx timestamp hang\n"); } static long wx_ptp_do_aux_work(struct ptp_clock_info *ptp) @@ -321,6 +363,9 @@ static long wx_ptp_do_aux_work(struct ptp_clock_info *ptp) struct wx *wx = container_of(ptp, struct wx, ptp_caps); int ts_done; + if (!test_bit(WX_STATE_PTP_RUNNING, wx->state)) + return HZ; + ts_done = wx_ptp_tx_hwtstamp_work(wx); wx_ptp_overflow_check(wx); @@ -836,6 +881,36 @@ void wx_ptp_stop(struct wx *wx) } EXPORT_SYMBOL(wx_ptp_stop); +void wx_ptp_quiesce(struct wx *wx) +{ + struct sk_buff *skb; + unsigned long flags; + + if (!test_and_clear_bit(WX_STATE_PTP_RUNNING, wx->state)) + return; + + clear_bit(WX_FLAG_PTP_PPS_ENABLED, wx->flags); + + if (wx->ptp_clock) + ptp_cancel_worker_sync(wx->ptp_clock); + + /* Drop a pending Tx timestamp request. Do not touch the registers + * here: quiesce runs during PCIe error recovery, where the device may + * already be gone and MMIO is not reliable. + */ + spin_lock_irqsave(&wx->ptp_tx_lock, flags); + skb = __wx_ptp_detach_tx_skb(wx); + spin_unlock_irqrestore(&wx->ptp_tx_lock, flags); + dev_kfree_skb_any(skb); + + if (wx->ptp_clock) { + ptp_clock_unregister(wx->ptp_clock); + wx->ptp_clock = NULL; + dev_info(&wx->pdev->dev, "removed PHC on %s\n", wx->netdev->name); + } +} +EXPORT_SYMBOL(wx_ptp_quiesce); + /** * wx_ptp_rx_hwtstamp - utility function which checks for RX time stamp * @wx: pointer to wx struct diff --git a/drivers/net/ethernet/wangxun/libwx/wx_ptp.h b/drivers/net/ethernet/wangxun/libwx/wx_ptp.h index 50db90a6e3ee..ad2f824875d5 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_ptp.h +++ b/drivers/net/ethernet/wangxun/libwx/wx_ptp.h @@ -10,6 +10,7 @@ void wx_ptp_reset(struct wx *wx); void wx_ptp_init(struct wx *wx); void wx_ptp_suspend(struct wx *wx); void wx_ptp_stop(struct wx *wx); +void wx_ptp_quiesce(struct wx *wx); void wx_ptp_rx_hwtstamp(struct wx *wx, struct sk_buff *skb); int wx_hwtstamp_get(struct net_device *dev, struct kernel_hwtstamp_config *cfg); diff --git a/drivers/net/ethernet/wangxun/libwx/wx_type.h b/drivers/net/ethernet/wangxun/libwx/wx_type.h index 0520288d18ab..e63a1666d8ed 100644 --- a/drivers/net/ethernet/wangxun/libwx/wx_type.h +++ b/drivers/net/ethernet/wangxun/libwx/wx_type.h @@ -1413,6 +1413,8 @@ struct wx { unsigned long last_overflow_check; unsigned long last_rx_ptp_check; unsigned long ptp_tx_start; + /* protects ptp_tx_skb, ptp_tx_start and the in-progress state bit */ + spinlock_t ptp_tx_lock; seqlock_t hw_tc_lock; /* seqlock for ptp */ struct cyclecounter hw_cc; struct timecounter hw_tc; diff --git a/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c b/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c index a84010828551..59a47532618c 100644 --- a/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c +++ b/drivers/net/ethernet/wangxun/txgbe/txgbe_fdir.c @@ -591,15 +591,24 @@ static void txgbe_fdir_filter_restore(struct wx *wx) queue = TXGBE_RDB_FDIR_DROP_QUEUE; } else { u32 ring = ethtool_get_flow_spec_ring(filter->action); + u8 vf = ethtool_get_flow_spec_ring_vf(filter->action); - if (ring >= wx->num_rx_queues) { + if (!vf && ring >= wx->num_rx_queues) { wx_err(wx, "FDIR restore failed, ring:%u\n", ring); continue; + } else if (vf && (vf > wx->num_vfs || + ring >= wx->num_rx_queues_per_pool)) { + wx_err(wx, "FDIR restore failed, vf:%u, ring:%u\n", + vf, ring); + continue; } /* Map the ring onto the absolute queue index */ - queue = wx->rx_ring[ring]->reg_idx; + if (!vf) + queue = wx->rx_ring[ring]->reg_idx; + else + queue = ((vf - 1) * wx->num_rx_queues_per_pool) + ring; } ret = txgbe_fdir_write_perfect_filter(wx, diff --git a/drivers/net/ethernet/wangxun/txgbe/txgbe_main.c b/drivers/net/ethernet/wangxun/txgbe/txgbe_main.c index c277863baf67..00123c0e3f22 100644 --- a/drivers/net/ethernet/wangxun/txgbe/txgbe_main.c +++ b/drivers/net/ethernet/wangxun/txgbe/txgbe_main.c @@ -93,12 +93,24 @@ static void txgbe_module_detection_subtask(struct wx *wx) { int err; + if (test_bit(WX_STATE_DOWN, wx->state) || + test_bit(WX_STATE_RESETTING, wx->state)) + return; + if (!test_and_clear_bit(WX_FLAG_NEED_MODULE_RESET, wx->flags)) return; /* wait for SFF module ready */ msleep(200); + /* Re-check state to avoid racing with down/reset paths. + * Module identification is deferred to the next up event, + * so it is safe to bail out here. + */ + if (test_bit(WX_STATE_DOWN, wx->state) || + test_bit(WX_STATE_RESETTING, wx->state)) + return; + err = txgbe_identify_module(wx); if (err == -ENODEV) set_bit(WX_FLAG_NEED_MODULE_RESET, wx->flags); @@ -106,6 +118,10 @@ static void txgbe_module_detection_subtask(struct wx *wx) static void txgbe_link_config_subtask(struct wx *wx) { + if (test_bit(WX_STATE_DOWN, wx->state) || + test_bit(WX_STATE_RESETTING, wx->state)) + return; + if (!test_and_clear_bit(WX_FLAG_NEED_LINK_CONFIG, wx->flags)) return; diff --git a/drivers/net/macsec.c b/drivers/net/macsec.c index ee0e2eb7dbc6..e4871cd76579 100644 --- a/drivers/net/macsec.c +++ b/drivers/net/macsec.c @@ -3539,6 +3539,22 @@ static int macsec_dev_init(struct net_device *dev) if (err) return err; + err = -ENOMEM; + macsec->stats = netdev_alloc_pcpu_stats(struct pcpu_secy_stats); + if (!macsec->stats) + goto destroy_gro_cells; + + macsec->secy.tx_sc.stats = + netdev_alloc_pcpu_stats(struct pcpu_tx_sc_stats); + if (!macsec->secy.tx_sc.stats) + goto free_secy_stats; + + macsec->secy.tx_sc.md_dst = metadata_dst_alloc(0, METADATA_MACSEC, + GFP_KERNEL); + if (!macsec->secy.tx_sc.md_dst) + goto free_tx_sc_stats; + macsec->secy.tx_sc.md_dst->u.macsec_info.sci = macsec->secy.sci; + macsec_inherit_tso_max(dev); dev->hw_features = real_dev->hw_features & MACSEC_OFFLOAD_FEATURES; @@ -3551,8 +3567,6 @@ static int macsec_dev_init(struct net_device *dev) macsec_set_head_tail_room(dev); - if (is_zero_ether_addr(dev->dev_addr)) - eth_hw_addr_inherit(dev, real_dev); if (is_zero_ether_addr(dev->broadcast)) memcpy(dev->broadcast, real_dev->broadcast, dev->addr_len); @@ -3560,6 +3574,14 @@ static int macsec_dev_init(struct net_device *dev) netdev_hold(real_dev, &macsec->dev_tracker, GFP_KERNEL); return 0; + +free_tx_sc_stats: + free_percpu(macsec->secy.tx_sc.stats); +free_secy_stats: + free_percpu(macsec->stats); +destroy_gro_cells: + gro_cells_destroy(&macsec->gro_cells); + return err; } static void macsec_dev_uninit(struct net_device *dev) @@ -4114,26 +4136,11 @@ static sci_t dev_to_sci(struct net_device *dev, __be16 port) return make_sci(dev->dev_addr, port); } -static int macsec_add_dev(struct net_device *dev, sci_t sci, u8 icv_len) +static void macsec_init_secy(struct net_device *dev, sci_t sci, u8 icv_len) { struct macsec_dev *macsec = macsec_priv(dev); struct macsec_secy *secy = &macsec->secy; - macsec->stats = netdev_alloc_pcpu_stats(struct pcpu_secy_stats); - if (!macsec->stats) - return -ENOMEM; - - secy->tx_sc.stats = netdev_alloc_pcpu_stats(struct pcpu_tx_sc_stats); - if (!secy->tx_sc.stats) - return -ENOMEM; - - secy->tx_sc.md_dst = metadata_dst_alloc(0, METADATA_MACSEC, GFP_KERNEL); - if (!secy->tx_sc.md_dst) - /* macsec and secy percpu stats will be freed when unregistering - * net_device in macsec_free_netdev() - */ - return -ENOMEM; - if (sci == MACSEC_UNDEF_SCI) sci = dev_to_sci(dev, MACSEC_PORT_ES); @@ -4147,15 +4154,12 @@ static int macsec_add_dev(struct net_device *dev, sci_t sci, u8 icv_len) secy->xpn = DEFAULT_XPN; secy->sci = sci; - secy->tx_sc.md_dst->u.macsec_info.sci = sci; secy->tx_sc.active = true; secy->tx_sc.encoding_sa = DEFAULT_ENCODING_SA; secy->tx_sc.encrypt = DEFAULT_ENCRYPT; secy->tx_sc.send_sci = DEFAULT_SEND_SCI; secy->tx_sc.end_station = false; secy->tx_sc.scb = false; - - return 0; } static struct lock_class_key macsec_netdev_addr_lock_key; @@ -4218,6 +4222,24 @@ static int macsec_newlink(struct net_device *dev, if (rx_handler && rx_handler != macsec_handle_frame) return -EBUSY; + if (is_zero_ether_addr(dev->dev_addr)) + eth_hw_addr_inherit(dev, real_dev); + + if (data && data[IFLA_MACSEC_SCI]) + sci = nla_get_sci(data[IFLA_MACSEC_SCI]); + else if (data && data[IFLA_MACSEC_PORT]) + sci = dev_to_sci(dev, nla_get_be16(data[IFLA_MACSEC_PORT])); + else + sci = dev_to_sci(dev, MACSEC_PORT_ES); + + /* Registration can notify listeners before returning. */ + macsec_init_secy(dev, sci, icv_len); + if (data) { + err = macsec_changelink_common(dev, data); + if (err) + return err; + } + err = register_netdevice(dev); if (err < 0) return err; @@ -4230,31 +4252,11 @@ static int macsec_newlink(struct net_device *dev, if (err < 0) goto unregister; - /* need to be already registered so that ->init has run and - * the MAC addr is set - */ - if (data && data[IFLA_MACSEC_SCI]) - sci = nla_get_sci(data[IFLA_MACSEC_SCI]); - else if (data && data[IFLA_MACSEC_PORT]) - sci = dev_to_sci(dev, nla_get_be16(data[IFLA_MACSEC_PORT])); - else - sci = dev_to_sci(dev, MACSEC_PORT_ES); - if (rx_handler && sci_exists(real_dev, sci)) { err = -EBUSY; goto unlink; } - err = macsec_add_dev(dev, sci, icv_len); - if (err) - goto unlink; - - if (data) { - err = macsec_changelink_common(dev, data); - if (err) - goto del_dev; - } - /* If h/w offloading is available, propagate to the device */ if (macsec_is_offloaded(macsec)) { const struct macsec_ops *ops; diff --git a/drivers/net/mdio/mdio-realtek-rtl9300.c b/drivers/net/mdio/mdio-realtek-rtl9300.c index afd52a1cd7f8..9ce2b7807532 100644 --- a/drivers/net/mdio/mdio-realtek-rtl9300.c +++ b/drivers/net/mdio/mdio-realtek-rtl9300.c @@ -88,6 +88,8 @@ #define RTL9310_SMI_INDRT_ACCESS_BC_PHYID_CTRL 0x0c14 #define RTL9310_BC_PORT_ID GENMASK(10, 5) #define RTL9310_SMI_INDRT_ACCESS_CTRL_1 0x0c04 +#define RTL9310_SMI_INDRT_EXT_PAGE GENMASK(8, 0) +#define RTL9310_SMI_INDRT_EXT_PAGE_NO_CHANGE 0x1ff #define RTL9310_SMI_INDRT_ACCESS_CTRL_2_LOW 0x0c08 #define RTL9310_SMI_INDRT_ACCESS_CTRL_2_HIGH 0x0c0c #define RTL9310_SMI_INDRT_ACCESS_CTRL_3 0x0c10 /* I/O fields flipped */ @@ -325,6 +327,8 @@ static int otto_emdio_9310_read_c22(struct mii_bus *bus, int port, int regnum, u .broadcast = FIELD_PREP(RTL9310_BC_PORT_ID, port), .c22_data = FIELD_PREP(RTL9310_PHY_CTRL_REG_ADDR, regnum) | FIELD_PREP(RTL9310_PHY_CTRL_MAIN_PAGE, RAW_PAGE(priv)), + .ext_page = FIELD_PREP(RTL9310_SMI_INDRT_EXT_PAGE, + RTL9310_SMI_INDRT_EXT_PAGE_NO_CHANGE), }; return otto_emdio_read_cmd(bus, RTL9310_PHY_CTRL_TYPE_C22, &cmd_data, @@ -337,6 +341,8 @@ static int otto_emdio_9310_write_c22(struct mii_bus *bus, int port, int regnum, struct otto_emdio_cmd_regs cmd_data = { .c22_data = FIELD_PREP(RTL9310_PHY_CTRL_REG_ADDR, regnum) | FIELD_PREP(RTL9310_PHY_CTRL_MAIN_PAGE, RAW_PAGE(priv)), + .ext_page = FIELD_PREP(RTL9310_SMI_INDRT_EXT_PAGE, + RTL9310_SMI_INDRT_EXT_PAGE_NO_CHANGE), .io_data = FIELD_PREP(RTL9310_PHY_CTRL_INDATA, value), .port_mask_high = (u32)(BIT_ULL(port) >> 32), .port_mask_low = (u32)(BIT_ULL(port)), diff --git a/drivers/net/ovpn/netlink.c b/drivers/net/ovpn/netlink.c index 4dad85294198..5432bc2eb8e8 100644 --- a/drivers/net/ovpn/netlink.c +++ b/drivers/net/ovpn/netlink.c @@ -100,6 +100,8 @@ static bool ovpn_nl_attr_sockaddr_remote(struct nlattr **attrs, struct sockaddr_in6 *sin6; struct sockaddr_in *sin; struct in6_addr *in6; + struct nlattr *scope; + u32 scope_id = 0; __be16 port = 0; __be32 *in; @@ -114,6 +116,9 @@ static bool ovpn_nl_attr_sockaddr_remote(struct nlattr **attrs, } else if (attrs[OVPN_A_PEER_REMOTE_IPV6]) { ss->ss_family = AF_INET6; in6 = nla_data(attrs[OVPN_A_PEER_REMOTE_IPV6]); + scope = attrs[OVPN_A_PEER_REMOTE_IPV6_SCOPE_ID]; + if (scope) + scope_id = nla_get_u32(scope); } else { return false; } @@ -126,6 +131,7 @@ static bool ovpn_nl_attr_sockaddr_remote(struct nlattr **attrs, if (!ipv6_addr_v4mapped(in6)) { sin6 = (struct sockaddr_in6 *)ss; sin6->sin6_port = port; + sin6->sin6_scope_id = scope_id; memcpy(&sin6->sin6_addr, in6, sizeof(*in6)); break; } @@ -179,6 +185,39 @@ static sa_family_t ovpn_nl_family_get(struct nlattr *addr4, return AF_UNSPEC; } +static int ovpn_nl_peer_check_vpn_addrs(const struct in_addr *addr4, + const struct in6_addr *addr6, + struct genl_info *info) +{ + int addr6_type; + + if (addr4->s_addr == htonl(INADDR_ANY) && ipv6_addr_any(addr6)) { + NL_SET_ERR_MSG_MOD(info->extack, + "at least one VPN IP must be configured in MP mode"); + return -EINVAL; + } + + if (ipv4_is_multicast(addr4->s_addr) || ipv4_is_lbcast(addr4->s_addr) || + ipv4_is_loopback(addr4->s_addr)) { + NL_SET_ERR_MSG_MOD(info->extack, + "VPN IPv4 address must be valid unicast or any"); + return -EADDRNOTAVAIL; + } + + if (!ipv6_addr_any(addr6)) { + addr6_type = ipv6_addr_type(addr6); + + if (!(addr6_type & IPV6_ADDR_UNICAST) || + (addr6_type & (IPV6_ADDR_LOOPBACK | IPV6_ADDR_COMPATv4))) { + NL_SET_ERR_MSG_MOD(info->extack, + "VPN IPv6 address must be valid unicast or any"); + return -EADDRNOTAVAIL; + } + } + + return 0; +} + static int ovpn_nl_peer_precheck(struct ovpn_priv *ovpn, struct genl_info *info, struct nlattr **attrs) @@ -346,8 +385,10 @@ static int ovpn_nl_peer_modify(struct ovpn_peer *peer, struct genl_info *info, int ovpn_nl_peer_new_doit(struct sk_buff *skb, struct genl_info *info) { - struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; + struct in_addr vpn_addr4 = { .s_addr = htonl(INADDR_ANY) }; + struct in6_addr vpn_addr6 = IN6ADDR_ANY_INIT; struct ovpn_priv *ovpn = info->user_ptr[0]; + struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; struct ovpn_socket *ovpn_sock; struct socket *sock = NULL; struct ovpn_peer *peer; @@ -371,11 +412,18 @@ int ovpn_nl_peer_new_doit(struct sk_buff *skb, struct genl_info *info) return -EINVAL; /* in MP mode VPN IPs are required for selecting the right peer */ - if (ovpn->mode == OVPN_MODE_MP && !attrs[OVPN_A_PEER_VPN_IPV4] && - !attrs[OVPN_A_PEER_VPN_IPV6]) { - NL_SET_ERR_MSG_FMT_MOD(info->extack, - "VPN IP must be provided in MP mode"); - return -EINVAL; + if (ovpn->mode == OVPN_MODE_MP) { + if (attrs[OVPN_A_PEER_VPN_IPV4]) + vpn_addr4.s_addr = + nla_get_in_addr(attrs[OVPN_A_PEER_VPN_IPV4]); + if (attrs[OVPN_A_PEER_VPN_IPV6]) + vpn_addr6 = + nla_get_in6_addr(attrs[OVPN_A_PEER_VPN_IPV6]); + + ret = ovpn_nl_peer_check_vpn_addrs(&vpn_addr4, &vpn_addr6, + info); + if (ret < 0) + return ret; } peer_id = nla_get_u32(attrs[OVPN_A_PEER_ID]); @@ -474,8 +522,10 @@ int ovpn_nl_peer_new_doit(struct sk_buff *skb, struct genl_info *info) int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info) { - struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; struct ovpn_priv *ovpn = info->user_ptr[0]; + struct nlattr *attrs[OVPN_A_PEER_MAX + 1]; + struct in6_addr vpn_addr6; + struct in_addr vpn_addr4; struct ovpn_socket *sock; struct ovpn_peer *peer; u32 peer_id; @@ -522,28 +572,58 @@ int ovpn_nl_peer_set_doit(struct sk_buff *skb, struct genl_info *info) rcu_read_unlock(); spin_lock_bh(&ovpn->lock); - ret = ovpn_nl_peer_modify(peer, info, attrs); - if (ret < 0) { - spin_unlock_bh(&ovpn->lock); - ovpn_peer_put(peer); - return ret; + + vpn_addr4 = peer->vpn_addrs.ipv4; + vpn_addr6 = peer->vpn_addrs.ipv6; + + /* reject peer with conflicting VPN address */ + if (attrs[OVPN_A_PEER_VPN_IPV4]) { + vpn_addr4.s_addr = nla_get_in_addr(attrs[OVPN_A_PEER_VPN_IPV4]); + if (ovpn_peer_vpn_addr_conflict4(ovpn, peer, &vpn_addr4)) + goto addr_conflict; + } + if (attrs[OVPN_A_PEER_VPN_IPV6]) { + vpn_addr6 = nla_get_in6_addr(attrs[OVPN_A_PEER_VPN_IPV6]); + if (ovpn_peer_vpn_addr_conflict6(ovpn, peer, &vpn_addr6)) + goto addr_conflict; } + /* in MP mode VPN IPs are required for selecting the right peer */ + if (ovpn->mode == OVPN_MODE_MP) { + ret = ovpn_nl_peer_check_vpn_addrs(&vpn_addr4, &vpn_addr6, + info); + if (ret < 0) + goto unlock; + } + + ret = ovpn_nl_peer_modify(peer, info, attrs); + if (ret < 0) + goto unlock; + /* ret == 1 means that VPN IPv4/6 has been modified and rehashing * is required */ - if (ret > 0) + if (ret > 0) { ovpn_peer_hash_vpn_ip(peer); + ret = 0; + } /* if the remote endpoint was updated, the by_transp_addr hash bucket * also needs to be refreshed, otherwise incoming packets from the new * remote address would fail the lockless lookup */ if (attrs[OVPN_A_PEER_REMOTE_IPV4] || attrs[OVPN_A_PEER_REMOTE_IPV6]) ovpn_peer_hash_transp_addr(peer); + +unlock: spin_unlock_bh(&ovpn->lock); ovpn_peer_put(peer); - return 0; + return ret; +addr_conflict: + NL_SET_ERR_MSG_FMT_MOD(info->extack, + "VPN IP is already assigned to another peer"); + ret = -EADDRINUSE; + goto unlock; } static int ovpn_nl_send_peer(struct sk_buff *skb, const struct genl_info *info, diff --git a/drivers/net/ovpn/peer.c b/drivers/net/ovpn/peer.c index c95656ca7c35..2067825bb5b6 100644 --- a/drivers/net/ovpn/peer.c +++ b/drivers/net/ovpn/peer.c @@ -113,6 +113,7 @@ struct ovpn_peer *ovpn_peer_new(struct ovpn_priv *ovpn, u32 id) RCU_INIT_POINTER(peer->bind, NULL); ovpn_crypto_state_init(&peer->crypto); spin_lock_init(&peer->lock); + seqcount_spinlock_init(&peer->route_key_seq, &peer->lock); kref_init(&peer->refcount); ovpn_peer_stats_init(&peer->vpn_stats); ovpn_peer_stats_init(&peer->link_stats); @@ -199,13 +200,12 @@ static void __ovpn_peer_hash_transp_addr(struct ovpn_peer *peer, */ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) { + const void *local_ip = NULL; struct sockaddr_storage ss; struct sockaddr_in6 *sa6; - bool reset_cache = false; struct sockaddr_in *sa; struct ovpn_bind *bind; - const void *local_ip; - size_t salen = 0; + bool floated = false; spin_lock_bh(&peer->lock); bind = rcu_dereference_protected(peer->bind, @@ -232,8 +232,7 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) .sin_addr.s_addr = ip_hdr(skb)->saddr, .sin_port = udp_hdr(skb)->source, }; - salen = sizeof(*sa); - reset_cache = true; + floated = true; break; } @@ -245,10 +244,12 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) netdev_name(peer->ovpn->dev), peer->id, &bind->local.ipv4.s_addr, &ip_hdr(skb)->daddr); - bind->local.ipv4.s_addr = ip_hdr(skb)->daddr; - reset_cache = true; + local_ip = &ip_hdr(skb)->daddr; + memcpy(&ss, &bind->remote, sizeof(struct sockaddr_in)); + break; } - break; + /* nothing changed */ + goto unlock; case htons(ETH_P_IPV6): /* float check */ if (unlikely(!ovpn_bind_skb_src_match(bind, skb))) { @@ -270,8 +271,7 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) ipv6_iface_scope_id(&ipv6_hdr(skb)->saddr, skb->skb_iif), }; - salen = sizeof(*sa6); - reset_cache = true; + floated = true; break; } @@ -284,26 +284,30 @@ void ovpn_peer_endpoints_update(struct ovpn_peer *peer, struct sk_buff *skb) netdev_name(peer->ovpn->dev), peer->id, &bind->local.ipv6, &ipv6_hdr(skb)->daddr); - bind->local.ipv6 = ipv6_hdr(skb)->daddr; - reset_cache = true; + local_ip = &ipv6_hdr(skb)->daddr; + memcpy(&ss, &bind->remote, sizeof(struct sockaddr_in6)); + break; } - break; + /* nothing changed */ + goto unlock; default: goto unlock; } - if (unlikely(reset_cache)) - dst_cache_reset(&peer->dst_cache); - - /* if the peer did not float, we can bail out now */ - if (likely(!salen)) - goto unlock; - if (unlikely(ovpn_peer_reset_sockaddr(peer, (struct sockaddr_storage *)&ss, local_ip) < 0)) goto unlock; + /* reset the cache only after a successful bind update to avoid useless + * cache misses on concurrent TX + */ + dst_cache_reset(&peer->dst_cache); + + /* if only the local address changed, bail out now */ + if (!floated) + goto unlock; + net_dbg_ratelimited("%s: peer %d floated to %pIScp", netdev_name(peer->ovpn->dev), peer->id, &ss); @@ -484,7 +488,7 @@ static struct ovpn_peer *ovpn_peer_get_by_vpn_addr4(struct ovpn_priv *ovpn, * Return: the peer if found or NULL otherwise */ static struct ovpn_peer *ovpn_peer_get_by_vpn_addr6(struct ovpn_priv *ovpn, - struct in6_addr *addr) + const struct in6_addr *addr) { struct hlist_nulls_head *nhead; struct hlist_nulls_node *ntmp; @@ -509,6 +513,64 @@ static struct ovpn_peer *ovpn_peer_get_by_vpn_addr6(struct ovpn_priv *ovpn, return NULL; } +/** + * ovpn_peer_vpn_addr_conflict4 - check if the VPN v4 address is already in use + * @ovpn: the openvpn instance to search + * @peer: peer being added or updated, or NULL + * @addr: VPN IPv4 address to check + * + * Check whether @addr is already assigned to another peer. @peer is ignored + * when found, allowing peer updates that keep an existing address. + * Unspecified addresses are ignored. + * + * Note: the caller must hold @ovpn->lock. + * + * Return: true on conflict, false otherwise. + */ +bool ovpn_peer_vpn_addr_conflict4(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in_addr *addr) +{ + struct ovpn_peer *tmp = NULL; + + lockdep_assert_held(&ovpn->lock); + + /* we don't hash INADDR_ANY, no conflict in that case */ + if (addr->s_addr != htonl(INADDR_ANY)) + tmp = ovpn_peer_get_by_vpn_addr4(ovpn, addr->s_addr); + + return tmp && tmp != peer; +} + +/** + * ovpn_peer_vpn_addr_conflict6 - check if the VPN v6 address is already in use + * @ovpn: the openvpn instance to search + * @peer: peer being added or updated, or NULL + * @addr: VPN IPv6 address to check + * + * Check whether @addr is already assigned to another peer. @peer is ignored + * when found, allowing peer updates that keep an existing address. + * Unspecified addresses are ignored. + * + * Note: the caller must hold @ovpn->lock. + * + * Return: true on conflict, false otherwise. + */ +bool ovpn_peer_vpn_addr_conflict6(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in6_addr *addr) +{ + struct ovpn_peer *tmp = NULL; + + lockdep_assert_held(&ovpn->lock); + + /* we don't hash ::, no conflict in that case */ + if (!ipv6_addr_any(addr)) + tmp = ovpn_peer_get_by_vpn_addr6(ovpn, addr); + + return tmp && tmp != peer; +} + /** * ovpn_peer_transp_match - check if sockaddr and peer binding match * @peer: the peer to get the binding from @@ -990,10 +1052,11 @@ void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer) if (hlist_unhashed(&peer->hash_entry_id)) return; - if (peer->vpn_addrs.ipv4.s_addr != htonl(INADDR_ANY)) { - /* remove potential old hashing */ - hlist_nulls_del_init_rcu(&peer->hash_entry_addr4); + /* remove potential old hashing */ + hlist_nulls_del_init_rcu(&peer->hash_entry_addr4); + hlist_nulls_del_init_rcu(&peer->hash_entry_addr6); + if (peer->vpn_addrs.ipv4.s_addr != htonl(INADDR_ANY)) { nhead = ovpn_get_hash_head(peer->ovpn->peers->by_vpn_addr4, &peer->vpn_addrs.ipv4, sizeof(peer->vpn_addrs.ipv4)); @@ -1001,9 +1064,6 @@ void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer) } if (!ipv6_addr_any(&peer->vpn_addrs.ipv6)) { - /* remove potential old hashing */ - hlist_nulls_del_init_rcu(&peer->hash_entry_addr6); - nhead = ovpn_get_hash_head(peer->ovpn->peers->by_vpn_addr6, &peer->vpn_addrs.ipv6, sizeof(peer->vpn_addrs.ipv6)); @@ -1038,6 +1098,13 @@ static int ovpn_peer_add_mp(struct ovpn_priv *ovpn, struct ovpn_peer *peer) goto out; } + /* reject peer with conflicting VPN address */ + if (ovpn_peer_vpn_addr_conflict4(ovpn, NULL, &peer->vpn_addrs.ipv4) || + ovpn_peer_vpn_addr_conflict6(ovpn, NULL, &peer->vpn_addrs.ipv6)) { + ret = -EADDRINUSE; + goto out; + } + bind = rcu_dereference_protected(peer->bind, true); /* peers connected via TCP have bind == NULL */ if (bind) { diff --git a/drivers/net/ovpn/peer.h b/drivers/net/ovpn/peer.h index dfa5c0037e02..1879bfb76992 100644 --- a/drivers/net/ovpn/peer.h +++ b/drivers/net/ovpn/peer.h @@ -10,6 +10,7 @@ #ifndef _NET_OVPN_OVPNPEER_H_ #define _NET_OVPN_OVPNPEER_H_ +#include #include #include @@ -17,6 +18,16 @@ #include "socket.h" #include "stats.h" +/** + * struct ovpn_route_key - route key used for the peer dst cache + * @mark: fwmark used for route lookup + * @sport: UDP source port used for route lookup + */ +struct ovpn_route_key { + u32 mark; + __be16 sport; +}; + /** * struct ovpn_peer - the main remote peer object * @ovpn: main openvpn instance this peer belongs to @@ -45,6 +56,8 @@ * @tcp.sk_cb.ops: pointer to the original prot_ops object (TCP only) * @crypto: the crypto configuration (ciphers, keys, etc..) * @dst_cache: cache for dst_entry used to send to peer + * @route_key: route key matching the current dst cache contents + * @route_key_seq: seqcount protecting lockless route_key reads * @bind: remote peer binding * @keepalive_interval: seconds after which a new keepalive should be sent * @keepalive_xmit_exp: future timestamp when next keepalive should be sent @@ -55,7 +68,7 @@ * @vpn_stats: per-peer in-VPN TX/RX stats * @link_stats: per-peer link/transport TX/RX stats * @delete_reason: why peer was deleted (i.e. timeout, transport error, ..) - * @lock: protects binding to peer (bind) and keepalive* fields + * @lock: protects binding to peer (bind), route_key and keepalive* fields * @refcount: reference counter * @rcu: used to free peer in an RCU safe way * @release_entry: entry for the socket release list @@ -99,6 +112,8 @@ struct ovpn_peer { } tcp; struct ovpn_crypto_state crypto; struct dst_cache dst_cache; + struct ovpn_route_key route_key; + seqcount_spinlock_t route_key_seq; struct ovpn_bind __rcu *bind; unsigned long keepalive_interval; unsigned long keepalive_xmit_exp; @@ -109,7 +124,7 @@ struct ovpn_peer { struct ovpn_peer_stats vpn_stats; struct ovpn_peer_stats link_stats; enum ovpn_del_peer_reason delete_reason; - spinlock_t lock; /* protects bind and keepalive* */ + spinlock_t lock; /* protects bind, route_key and keepalive* */ struct kref refcount; struct rcu_head rcu; struct llist_node release_entry; @@ -149,6 +164,12 @@ struct ovpn_peer *ovpn_peer_get_by_transp_addr(struct ovpn_priv *ovpn, struct ovpn_peer *ovpn_peer_get_by_id(struct ovpn_priv *ovpn, u32 peer_id); struct ovpn_peer *ovpn_peer_get_by_dst(struct ovpn_priv *ovpn, struct sk_buff *skb); +bool ovpn_peer_vpn_addr_conflict4(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in_addr *addr); +bool ovpn_peer_vpn_addr_conflict6(struct ovpn_priv *ovpn, + const struct ovpn_peer *peer, + const struct in6_addr *addr); void ovpn_peer_hash_vpn_ip(struct ovpn_peer *peer); void ovpn_peer_hash_transp_addr(struct ovpn_peer *peer); bool ovpn_peer_check_by_src(struct ovpn_priv *ovpn, struct sk_buff *skb, diff --git a/drivers/net/ovpn/udp.c b/drivers/net/ovpn/udp.c index 7f69e8890b5b..055cdb1bee13 100644 --- a/drivers/net/ovpn/udp.c +++ b/drivers/net/ovpn/udp.c @@ -131,6 +131,77 @@ static int ovpn_udp_encap_recv(struct sock *sk, struct sk_buff *skb) return 0; } +static bool ovpn_route_key_equal(const struct ovpn_route_key *a, + const struct ovpn_route_key *b) +{ + return a->mark == b->mark && a->sport == b->sport; +} + +/** + * ovpn_dst_cache_check_key - reset peer dst cache after key changes + * @peer: the peer owning the dst cache + * @cache: the cache that might need to be reset + * @key: the route key for the packet being transmitted + * + * Reset the peer dst cache if it was populated for a different route key. + */ +static void ovpn_dst_cache_check_key(struct ovpn_peer *peer, + struct dst_cache *cache, + const struct ovpn_route_key *key) +{ + struct ovpn_route_key old_key; + unsigned int seq; + + /* snapshot the saved key before deciding whether the cache matches */ + do { + seq = read_seqcount_begin(&peer->route_key_seq); + old_key = peer->route_key; + } while (read_seqcount_retry(&peer->route_key_seq, seq)); + + /* nothing changed: the current cache can be reused */ + if (likely(ovpn_route_key_equal(&old_key, key))) + return; + + /* recheck under lock because another path may have updated the key */ + spin_lock_bh(&peer->lock); + if (!ovpn_route_key_equal(&peer->route_key, key)) { + write_seqcount_begin(&peer->route_key_seq); + peer->route_key = *key; + dst_cache_reset(cache); + write_seqcount_end(&peer->route_key_seq); + } + spin_unlock_bh(&peer->lock); +} + +/** + * ovpn_dst_cache_current - check whether a route lookup matches peer state + * @peer: the peer owning the bind and dst cache + * @bind: the RCU bind used for the route lookup + * @key: the route key used for the route lookup + * + * Check that @bind is still the current peer bind and that @key still matches + * the peer route key. The caller must hold @peer->lock. The TX path keeps + * @bind inside an RCU read-side critical section, so pointer identity is enough + * to detect whether the bind was replaced while the route lookup was running. + * + * Return: true if the lookup result still matches the current peer state and + * may update the dst cache or replace the bind. + */ +static bool ovpn_dst_cache_current(const struct ovpn_peer *peer, + const struct ovpn_bind *bind, + const struct ovpn_route_key *key) +{ + const struct ovpn_bind *curr_bind; + + lockdep_assert_held(&peer->lock); + + curr_bind = rcu_dereference_protected(peer->bind, + lockdep_is_held(&peer->lock)); + + return curr_bind == bind && + ovpn_route_key_equal(key, &peer->route_key); +} + /** * ovpn_udp4_output - send IPv4 packet over udp socket * @peer: the destination peer @@ -138,21 +209,26 @@ static int ovpn_udp_encap_recv(struct sock *sk, struct sk_buff *skb) * @cache: dst cache * @sk: the socket to send the packet over * @skb: the packet to send + * @key: the route key snapshot used for cache validation and flow lookup * * Return: 0 on success or a negative error code otherwise */ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, struct dst_cache *cache, struct sock *sk, - struct sk_buff *skb) + struct sk_buff *skb, + const struct ovpn_route_key *key) { + struct sockaddr_storage remote; + struct in_addr local = {}; + bool reset_local = false; struct rtable *rt; struct flowi4 fl = { .saddr = bind->local.ipv4.s_addr, .daddr = bind->remote.in4.sin_addr.s_addr, - .fl4_sport = inet_sk(sk)->inet_sport, + .fl4_sport = key->sport, .fl4_dport = bind->remote.in4.sin_port, .flowi4_proto = sk->sk_protocol, - .flowi4_mark = sk->sk_mark, + .flowi4_mark = key->mark, }; int ret; @@ -161,26 +237,19 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, if (rt) goto transmit; - if (unlikely(!inet_confirm_addr(sock_net(sk), NULL, 0, fl.saddr, - RT_SCOPE_HOST))) { - /* we may end up here when the cached address is not usable - * anymore. In this case we reset address/cache and perform a - * new look up + if (fl.saddr && unlikely(!inet_confirm_addr(sock_net(sk), NULL, 0, + fl.saddr, RT_SCOPE_HOST))) { + /* The learned local address is not usable anymore. + * Retry with source address autoselection. */ fl.saddr = 0; - spin_lock_bh(&peer->lock); - bind->local.ipv4.s_addr = 0; - spin_unlock_bh(&peer->lock); - dst_cache_reset(cache); + reset_local = true; } rt = ip_route_output_flow(sock_net(sk), &fl, sk); if (IS_ERR(rt) && PTR_ERR(rt) == -EINVAL) { fl.saddr = 0; - spin_lock_bh(&peer->lock); - bind->local.ipv4.s_addr = 0; - spin_unlock_bh(&peer->lock); - dst_cache_reset(cache); + reset_local = true; rt = ip_route_output_flow(sock_net(sk), &fl, sk); } @@ -193,7 +262,30 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, ret); goto err; } - dst_cache_set_ip4(cache, &rt->dst, fl.saddr); + + /* avoid storing a stale cache or local address */ + spin_lock_bh(&peer->lock); + if (likely(ovpn_dst_cache_current(peer, bind, key))) { + if (!reset_local) { + dst_cache_set_ip4(cache, &rt->dst, fl.saddr); + spin_unlock_bh(&peer->lock); + goto transmit; + } + + /* invalidate per-CPU dst entries that may still carry + * the stale source + */ + dst_cache_reset(cache); + + /* preserve the current remote */ + memcpy(&remote, &bind->remote, sizeof(struct sockaddr_in)); + /* The current packet already has a valid wildcard-source route. + * If replacing the bind fails, leave the stale local in place; + * a later cache miss will retry the repair. + */ + ovpn_peer_reset_sockaddr(peer, &remote, &local); + } + spin_unlock_bh(&peer->lock); transmit: udp_tunnel_xmit_skb(rt, sk, skb, fl.saddr, fl.daddr, 0, @@ -213,23 +305,28 @@ static int ovpn_udp4_output(struct ovpn_peer *peer, struct ovpn_bind *bind, * @cache: dst cache * @sk: the socket to send the packet over * @skb: the packet to send + * @key: the route key snapshot used for cache validation and flow lookup * * Return: 0 on success or a negative error code otherwise */ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, struct dst_cache *cache, struct sock *sk, - struct sk_buff *skb) + struct sk_buff *skb, + const struct ovpn_route_key *key) { + struct in6_addr local = in6addr_any; + struct sockaddr_storage remote; + bool reset_local = false; struct dst_entry *dst; int ret; struct flowi6 fl = { .saddr = bind->local.ipv6, .daddr = bind->remote.in6.sin6_addr, - .fl6_sport = inet_sk(sk)->inet_sport, + .fl6_sport = key->sport, .fl6_dport = bind->remote.in6.sin6_port, .flowi6_proto = sk->sk_protocol, - .flowi6_mark = sk->sk_mark, + .flowi6_mark = key->mark, .flowi6_oif = bind->remote.in6.sin6_scope_id, }; @@ -238,16 +335,13 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, if (dst) goto transmit; - if (unlikely(!ipv6_chk_addr(sock_net(sk), &fl.saddr, NULL, 0))) { - /* we may end up here when the cached address is not usable - * anymore. In this case we reset address/cache and perform a - * new look up + if (!ipv6_addr_any(&fl.saddr) && + unlikely(!ipv6_chk_addr(sock_net(sk), &fl.saddr, NULL, 0))) { + /* The learned local address is not usable anymore. + * Retry with source address autoselection. */ fl.saddr = in6addr_any; - spin_lock_bh(&peer->lock); - bind->local.ipv6 = in6addr_any; - spin_unlock_bh(&peer->lock); - dst_cache_reset(cache); + reset_local = true; } dst = ip6_dst_lookup_flow(sock_net(sk), sk, &fl, NULL); @@ -258,7 +352,30 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, &bind->remote.in6, ret); goto err; } - dst_cache_set_ip6(cache, dst, &fl.saddr); + + /* avoid storing a stale cache or local address */ + spin_lock_bh(&peer->lock); + if (likely(ovpn_dst_cache_current(peer, bind, key))) { + if (!reset_local) { + dst_cache_set_ip6(cache, dst, &fl.saddr); + spin_unlock_bh(&peer->lock); + goto transmit; + } + + /* invalidate per-CPU dst entries that may still carry + * the stale source + */ + dst_cache_reset(cache); + + /* preserve the current remote */ + memcpy(&remote, &bind->remote, sizeof(struct sockaddr_in6)); + /* The current packet already has a valid wildcard-source route. + * If replacing the bind fails, leave the stale local in place; + * a later cache miss will retry the repair. + */ + ovpn_peer_reset_sockaddr(peer, &remote, &local); + } + spin_unlock_bh(&peer->lock); transmit: /* user IPv6 packets may be larger than the transport interface @@ -287,6 +404,7 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, * @cache: dst cache * @sk: the socket to send the packet over * @skb: the packet to send + * @key: route key snapshot used for cache validation and flow lookup * * rcu_read_lock should be held on entry. * On return, the skb is consumed. @@ -294,7 +412,8 @@ static int ovpn_udp6_output(struct ovpn_peer *peer, struct ovpn_bind *bind, * Return: 0 on success or a negative error code otherwise */ static int ovpn_udp_output(struct ovpn_peer *peer, struct dst_cache *cache, - struct sock *sk, struct sk_buff *skb) + struct sock *sk, struct sk_buff *skb, + struct ovpn_route_key *key) { struct ovpn_bind *bind; int ret; @@ -314,11 +433,11 @@ static int ovpn_udp_output(struct ovpn_peer *peer, struct dst_cache *cache, switch (bind->remote.in4.sin_family) { case AF_INET: - ret = ovpn_udp4_output(peer, bind, cache, sk, skb); + ret = ovpn_udp4_output(peer, bind, cache, sk, skb, key); break; #if IS_ENABLED(CONFIG_IPV6) case AF_INET6: - ret = ovpn_udp6_output(peer, bind, cache, sk, skb); + ret = ovpn_udp6_output(peer, bind, cache, sk, skb, key); break; #endif default: @@ -340,15 +459,21 @@ static int ovpn_udp_output(struct ovpn_peer *peer, struct dst_cache *cache, void ovpn_udp_send_skb(struct ovpn_peer *peer, struct sock *sk, struct sk_buff *skb) { + struct ovpn_route_key key = { + .mark = READ_ONCE(sk->sk_mark), + .sport = READ_ONCE(inet_sk(sk)->inet_sport), + }; int ret; skb->dev = peer->ovpn->dev; - skb->mark = READ_ONCE(sk->sk_mark); + skb->mark = key.mark; /* no checksum performed at this layer */ skb->ip_summed = CHECKSUM_NONE; + ovpn_dst_cache_check_key(peer, &peer->dst_cache, &key); + /* crypto layer -> transport (UDP) */ - ret = ovpn_udp_output(peer, &peer->dst_cache, sk, skb); + ret = ovpn_udp_output(peer, &peer->dst_cache, sk, skb, &key); if (unlikely(ret < 0)) kfree_skb(skb); } diff --git a/drivers/net/pcs/pcs-rzn1-miic.c b/drivers/net/pcs/pcs-rzn1-miic.c index 2b72fa98ddf1..cb74861e823c 100644 --- a/drivers/net/pcs/pcs-rzn1-miic.c +++ b/drivers/net/pcs/pcs-rzn1-miic.c @@ -683,7 +683,8 @@ static int miic_parse_dt(struct miic *miic, u32 *mode_cfg) if (!dt_val) return -ENOMEM; - memset(dt_val, MIIC_MODCTRL_CONF_NONE, sizeof(*dt_val)); + memset(dt_val, MIIC_MODCTRL_CONF_NONE, + sizeof(*dt_val) * miic->of_data->conf_conv_count); if (of_property_read_u32(np, "renesas,miic-switch-portin", &conf) == 0) dt_val[0] = conf; diff --git a/drivers/net/pcs/pcs-xpcs.c b/drivers/net/pcs/pcs-xpcs.c index 0337e2bcc012..b415b93d77c1 100644 --- a/drivers/net/pcs/pcs-xpcs.c +++ b/drivers/net/pcs/pcs-xpcs.c @@ -1545,8 +1545,10 @@ static int xpcs_init_clks(struct dw_xpcs *xpcs) return dev_err_probe(dev, ret, "Failed to get clocks\n"); ret = clk_bulk_prepare_enable(DW_XPCS_NUM_CLKS, xpcs->clks); - if (ret) + if (ret) { + clk_bulk_put(DW_XPCS_NUM_CLKS, xpcs->clks); return dev_err_probe(dev, ret, "Failed to enable clocks\n"); + } return 0; } diff --git a/drivers/net/phy/intel-xway.c b/drivers/net/phy/intel-xway.c index afbcec711744..3cee31bb931f 100644 --- a/drivers/net/phy/intel-xway.c +++ b/drivers/net/phy/intel-xway.c @@ -16,6 +16,11 @@ #define XWAY_MDIO_ISTAT 0x1A /* interrupt status */ #define XWAY_MDIO_LED 0x1B /* led control */ +#define XWAY_MDIO_GCTRL_TM_MASK GENMASK(15, 13) +#define XWAY_MDIO_GCTRL_TM(mode) FIELD_PREP(XWAY_MDIO_GCTRL_TM_MASK, (mode)) +#define XWAY_MDIO_GCTRL_TM_NOP XWAY_MDIO_GCTRL_TM(0) /* Normal operation */ +#define XWAY_MDIO_GCTRL_TM_CDIAG XWAY_MDIO_GCTRL_TM(6) /* Cable diagnostics */ + #define XWAY_MDIO_ERRCNT_SEL GENMASK(11, 8) #define XWAY_MDIO_ERRCNT_COUNT GENMASK(7, 0) #define XWAY_MDIO_ERRCNT_SEL_RXERR 0 @@ -326,6 +331,28 @@ static int xway_gphy_probe(struct phy_device *phydev) return 0; } +static int xway_11g_int_config_init(struct phy_device *phydev) +{ + int err; + + /* An issue has been sporadically observed after device power-on on the + * first link-up attempt in 100BASE-TX mode resulting in either the + * link-up taking a long time, or failing to link-up altogether. + * + * Workaround: + * After power-on, enable Cable Diagnostic Mode for all ports and + * disable it. + */ + err = phy_modify(phydev, MII_CTRL1000, XWAY_MDIO_GCTRL_TM_MASK, XWAY_MDIO_GCTRL_TM_CDIAG); + if (err) + return err; + err = phy_modify(phydev, MII_CTRL1000, XWAY_MDIO_GCTRL_TM_MASK, XWAY_MDIO_GCTRL_TM_NOP); + if (err) + return err; + + return xway_gphy_config_init(phydev); +} + static int xway_gphy14_config_aneg(struct phy_device *phydev) { int reg, err; @@ -735,7 +762,7 @@ static struct phy_driver xway_gphy[] = { .phy_id_mask = 0xffffffff, .name = "Intel XWAY PHY11G (xRX v1.2 integrated)", /* PHY_GBIT_FEATURES */ - .config_init = xway_gphy_config_init, + .config_init = xway_11g_int_config_init, .probe = xway_gphy_probe, .handle_interrupt = xway_gphy_handle_interrupt, .config_intr = xway_gphy_config_intr, diff --git a/drivers/net/phy/micrel.c b/drivers/net/phy/micrel.c index 55df5efcfc86..cde60498d054 100644 --- a/drivers/net/phy/micrel.c +++ b/drivers/net/phy/micrel.c @@ -6285,6 +6285,7 @@ static int lanphy_write_reg_data(struct phy_device *phydev, data->val); if (ret) break; + data++; } return ret; diff --git a/drivers/net/phy/phylink.c b/drivers/net/phy/phylink.c index 27fd9aa2027a..3477e1104a43 100644 --- a/drivers/net/phy/phylink.c +++ b/drivers/net/phy/phylink.c @@ -2131,7 +2131,6 @@ static int phylink_bringup_phy(struct phylink *pl, struct phy_device *phy, mutex_lock(&pl->phydev_mutex); mutex_lock(&phy->lock); mutex_lock(&pl->state_mutex); - pl->phydev = phy; pl->phy_state.interface = interface; pl->phy_state.pause = MLO_PAUSE_NONE; pl->phy_state.speed = SPEED_UNKNOWN; @@ -2198,10 +2197,25 @@ static int phylink_bringup_phy(struct phylink *pl, struct phy_device *phy, ret = 0; } - if (ret == 0 && phy_interrupt_is_valid(phy)) + if (ret) + return ret; + + /* Nothing below can fail, so the PHY can be recorded now. Doing it + * here rather than above keeps a failed bringup from leaving + * pl->phydev pointing at a PHY the caller is about to detach. + */ + mutex_lock(&pl->phydev_mutex); + mutex_lock(&phy->lock); + mutex_lock(&pl->state_mutex); + pl->phydev = phy; + mutex_unlock(&pl->state_mutex); + mutex_unlock(&phy->lock); + mutex_unlock(&pl->phydev_mutex); + + if (phy_interrupt_is_valid(phy)) phy_request_interrupt(phy); - return ret; + return 0; } static int phylink_attach_phy(struct phylink *pl, struct phy_device *phy, diff --git a/drivers/net/usb/catc.c b/drivers/net/usb/catc.c index 96e82f94edcf..39b678f175dd 100644 --- a/drivers/net/usb/catc.c +++ b/drivers/net/usb/catc.c @@ -233,17 +233,26 @@ static void catc_rx_done(struct urb *urb) } do { - if(!catc->is_f5u011) { - pkt_len = le16_to_cpup((__le16*)pkt_start); - if (pkt_len > urb->actual_length) { + int remaining = urb->actual_length - + (pkt_start - (u8 *)urb->transfer_buffer); + + if (!catc->is_f5u011) { + if (remaining < pkt_offset) { catc->netdev->stats.rx_length_errors++; catc->netdev->stats.rx_errors++; break; } + pkt_len = le16_to_cpup((__le16 *)pkt_start); } else { pkt_len = urb->actual_length; } + if (pkt_len < ETH_HLEN || pkt_len + pkt_offset > remaining) { + catc->netdev->stats.rx_length_errors++; + catc->netdev->stats.rx_errors++; + break; + } + if (!(skb = dev_alloc_skb(pkt_len))) return; diff --git a/drivers/net/usb/cdc_mbim.c b/drivers/net/usb/cdc_mbim.c index 877fb0ed7d3d..a7010a0664c7 100644 --- a/drivers/net/usb/cdc_mbim.c +++ b/drivers/net/usb/cdc_mbim.c @@ -635,6 +635,11 @@ static const struct usb_device_id mbim_devs[] = { .driver_info = (unsigned long)&cdc_mbim_info, }, + /* MeiG Smart SRM821 ZLP conformance */ + { USB_DEVICE_AND_INTERFACE_INFO(0x2dee, 0x4d53, USB_CLASS_COMM, USB_CDC_SUBCLASS_MBIM, USB_CDC_PROTO_NONE), + .driver_info = (unsigned long)&cdc_mbim_info, + }, + /* Some Huawei devices, ME906s-158 (12d1:15c1) and E3372 * (12d1:157d), are known to fail unless the NDP is placed * after the IP packets. Applying the quirk to all Huawei diff --git a/drivers/net/usb/lan78xx.c b/drivers/net/usb/lan78xx.c index cb782d81d84f..5655941f1478 100644 --- a/drivers/net/usb/lan78xx.c +++ b/drivers/net/usb/lan78xx.c @@ -5239,10 +5239,12 @@ static bool lan78xx_submit_deferred_urbs(struct lan78xx_net *dev) !netif_carrier_ok(dev->net) || pipe_halted) { lan78xx_release_tx_buf(dev, skb); + usb_put_urb(urb); continue; } ret = usb_submit_urb(urb, GFP_ATOMIC); + usb_put_urb(urb); if (ret == 0) { netif_trans_update(dev->net); diff --git a/drivers/net/usb/sr9700.c b/drivers/net/usb/sr9700.c index 937e6fef3ac6..50981a28376a 100644 --- a/drivers/net/usb/sr9700.c +++ b/drivers/net/usb/sr9700.c @@ -355,7 +355,8 @@ static int sr9700_rx_fixup(struct usbnet *dev, struct sk_buff *skb) /* ignore the CRC length */ len = (skb->data[1] | (skb->data[2] << 8)) - 4; - if (len > ETH_FRAME_LEN || len > skb->len || len < 0) + if (len > ETH_FRAME_LEN || len < 0 || + len > skb->len - SR_RX_OVERHEAD) return 0; /* the last packet of current skb */ diff --git a/drivers/net/veth.c b/drivers/net/veth.c index 6ab84c837a33..2780b2dd073d 100644 --- a/drivers/net/veth.c +++ b/drivers/net/veth.c @@ -1053,6 +1053,7 @@ static int __veth_napi_enable_range(struct net_device *dev, int start, int end) for (i = start; i < end; i++) { struct veth_rq *rq = &priv->rq[i]; + rcu_assign_pointer(rq->xdp_prog, priv->_xdp_prog); napi_enable(&rq->xdp_napi); rcu_assign_pointer(priv->rq[i].napi, &priv->rq[i].xdp_napi); } @@ -1087,6 +1088,7 @@ static void veth_napi_del_range(struct net_device *dev, int start, int end) rcu_assign_pointer(priv->rq[i].napi, NULL); napi_disable(&rq->xdp_napi); + rcu_assign_pointer(rq->xdp_prog, NULL); __netif_napi_del(&rq->xdp_napi); } synchronize_net(); diff --git a/drivers/net/virtio_net.c b/drivers/net/virtio_net.c index e34c52d059d3..bf82ef9874ab 100644 --- a/drivers/net/virtio_net.c +++ b/drivers/net/virtio_net.c @@ -3349,6 +3349,14 @@ static netdev_tx_t start_xmit(struct sk_buff *skb, struct net_device *dev) else virtqueue_disable_cb(sq->vq); + if (!use_napi && + unlikely(skb_orphan_frags(skb, GFP_ATOMIC))) { + DEV_STATS_INC(dev, tx_dropped); + dev_kfree_skb_any(skb); + kick = !xmit_more || netif_xmit_stopped(txq); + goto kick_vq; + } + /* timestamp packet in software */ skb_tx_timestamp(skb); @@ -3381,6 +3389,7 @@ static netdev_tx_t start_xmit(struct sk_buff *skb, struct net_device *dev) kick = use_napi ? __netdev_tx_sent_queue(txq, skb->len, xmit_more) : !xmit_more || netif_xmit_stopped(txq); +kick_vq: if (kick) { if (virtqueue_kick_prepare(sq->vq) && virtqueue_notify(sq->vq)) { u64_stats_update_begin(&sq->stats.syncp); diff --git a/drivers/net/vrf.c b/drivers/net/vrf.c index 46209917ae4d..bae0b69cf894 100644 --- a/drivers/net/vrf.c +++ b/drivers/net/vrf.c @@ -1175,8 +1175,6 @@ static int vrf_prepare_mac_header(struct sk_buff *skb, skb->protocol = eth->h_proto; skb->pkt_type = PACKET_HOST; - skb_postpush_rcsum(skb, skb->data, ETH_HLEN); - skb_pull_inline(skb, ETH_HLEN); return 0; diff --git a/drivers/net/vxlan/vxlan_core.c b/drivers/net/vxlan/vxlan_core.c index e045e1ee9e59..c4bcaadba471 100644 --- a/drivers/net/vxlan/vxlan_core.c +++ b/drivers/net/vxlan/vxlan_core.c @@ -1958,13 +1958,15 @@ static struct sk_buff *vxlan_na_create(struct sk_buff *request, struct ipv6hdr *pip6; u8 *daddr; int na_olen = 8; /* opt hdr + ETH_ALEN for target */ + int headroom; int ns_olen; int i, len; if (dev == NULL || !pskb_may_pull(request, request->len)) return NULL; - len = LL_RESERVED_SPACE(dev) + sizeof(struct ipv6hdr) + + headroom = LL_RESERVED_SPACE(dev); + len = headroom + sizeof(struct ipv6hdr) + sizeof(*na) + na_olen + dev->needed_tailroom; reply = alloc_skb(len, GFP_ATOMIC); if (reply == NULL) @@ -1972,7 +1974,7 @@ static struct sk_buff *vxlan_na_create(struct sk_buff *request, reply->protocol = htons(ETH_P_IPV6); reply->dev = dev; - skb_reserve(reply, LL_RESERVED_SPACE(request->dev)); + skb_reserve(reply, headroom); skb_push(reply, sizeof(struct ethhdr)); skb_reset_mac_header(reply); diff --git a/drivers/nfc/microread/microread.c b/drivers/nfc/microread/microread.c index dfa2490db545..2bfafa94e83d 100644 --- a/drivers/nfc/microread/microread.c +++ b/drivers/nfc/microread/microread.c @@ -251,9 +251,9 @@ static int microread_start_poll(struct nfc_hci_dev *hdev, param[1] |= (1 << 1); if ((im_protocols | tm_protocols) & NFC_PROTO_NFC_DEP_MASK) { - hdev->gb = nfc_get_local_general_bytes(hdev->ndev, - &hdev->gb_len); - if (hdev->gb == NULL || hdev->gb_len == 0) { + nfc_get_local_general_bytes(hdev->ndev, hdev->gb, + sizeof(hdev->gb), &hdev->gb_len); + if (hdev->gb_len == 0) { im_protocols &= ~NFC_PROTO_NFC_DEP_MASK; tm_protocols &= ~NFC_PROTO_NFC_DEP_MASK; } diff --git a/drivers/nfc/nfcmrvl/fw_dnld.c b/drivers/nfc/nfcmrvl/fw_dnld.c index 2b8f401d8fd7..8b9d5257320d 100644 --- a/drivers/nfc/nfcmrvl/fw_dnld.c +++ b/drivers/nfc/nfcmrvl/fw_dnld.c @@ -263,9 +263,14 @@ static int process_state_fw_dnld(struct nfcmrvl_private *priv, * B8..N: payload */ - /* Remove NCI HDR */ - skb_pull(skb, 3); - if (skb->data[0] != HELPER_CMD_PACKET_FORMAT || skb->len != 5) { + if (skb->len != NCI_DATA_HDR_SIZE + 5) { + nfc_err(priv->dev, "bad command"); + return -EINVAL; + } + + /* Remove NCI header */ + skb_pull(skb, NCI_DATA_HDR_SIZE); + if (skb->data[0] != HELPER_CMD_PACKET_FORMAT) { nfc_err(priv->dev, "bad command"); return -EINVAL; } diff --git a/drivers/nfc/pn533/pn533.c b/drivers/nfc/pn533/pn533.c index 76081a99b450..572f0f8737c1 100644 --- a/drivers/nfc/pn533/pn533.c +++ b/drivers/nfc/pn533/pn533.c @@ -1355,10 +1355,11 @@ static int pn533_poll_dep(struct nfc_dev *nfc_dev) u8 *next, nfcid3[NFC_NFCID3_MAXSIZE]; u8 passive_data[PASSIVE_DATA_LEN] = {0x00, 0xff, 0xff, 0x00, 0x3}; - if (!dev->gb) { - dev->gb = nfc_get_local_general_bytes(nfc_dev, &dev->gb_len); - - if (!dev->gb || !dev->gb_len) { + if (!dev->gb_len) { + nfc_get_local_general_bytes(nfc_dev, dev->gb, + sizeof(dev->gb), + &dev->gb_len); + if (!dev->gb_len) { dev->poll_dep = 0; queue_work(dev->wq, &dev->rf_work); } @@ -1656,8 +1657,9 @@ static int pn533_start_poll(struct nfc_dev *nfc_dev, } if (tm_protocols) { - dev->gb = nfc_get_local_general_bytes(nfc_dev, &dev->gb_len); - if (dev->gb == NULL) + nfc_get_local_general_bytes(nfc_dev, dev->gb, + sizeof(dev->gb), &dev->gb_len); + if (dev->gb_len == 0) tm_protocols = 0; } diff --git a/drivers/nfc/pn533/pn533.h b/drivers/nfc/pn533/pn533.h index 09e35b8693f5..5ab668e05121 100644 --- a/drivers/nfc/pn533/pn533.h +++ b/drivers/nfc/pn533/pn533.h @@ -6,6 +6,8 @@ * Copyright (C) 2012-2013 Tieto Poland */ +#include + #define PN533_DEVICE_STD 0x1 #define PN533_DEVICE_PASORI 0x2 #define PN533_DEVICE_ACR122U 0x3 @@ -166,7 +168,7 @@ struct pn533 { struct timer_list listen_timer; int cancel_listen; - u8 *gb; + u8 gb[NFC_MAX_GT_LEN]; size_t gb_len; u8 tgt_available_prots; diff --git a/drivers/nfc/pn533/usb.c b/drivers/nfc/pn533/usb.c index efb07f944fce..972eaac09e59 100644 --- a/drivers/nfc/pn533/usb.c +++ b/drivers/nfc/pn533/usb.c @@ -319,7 +319,9 @@ static bool pn533_acr122_is_rx_frame_valid(void *_frame, struct pn533 *dev) if (frame->ccid.type != 0x83) return false; - if (!frame->ccid.datalen) + if (frame->ccid.datalen < 2 || + frame->ccid.datalen > PN533_ACR122_FRAME_MAX_PAYLOAD_LEN + + PN533_ACR122_RX_FRAME_TAIL_LEN) return false; if (frame->data[frame->ccid.datalen - 2] == 0x63) diff --git a/drivers/nfc/pn544/pn544.c b/drivers/nfc/pn544/pn544.c index 9d0a16ac465e..c4fa70e45c14 100644 --- a/drivers/nfc/pn544/pn544.c +++ b/drivers/nfc/pn544/pn544.c @@ -377,10 +377,9 @@ static int pn544_hci_start_poll(struct nfc_hci_dev *hdev, return r; if ((im_protocols | tm_protocols) & NFC_PROTO_NFC_DEP_MASK) { - hdev->gb = nfc_get_local_general_bytes(hdev->ndev, - &hdev->gb_len); - pr_debug("generate local bytes %p\n", hdev->gb); - if (hdev->gb == NULL || hdev->gb_len == 0) { + nfc_get_local_general_bytes(hdev->ndev, hdev->gb, + sizeof(hdev->gb), &hdev->gb_len); + if (hdev->gb_len == 0) { im_protocols &= ~NFC_PROTO_NFC_DEP_MASK; tm_protocols &= ~NFC_PROTO_NFC_DEP_MASK; } diff --git a/drivers/nfc/port100.c b/drivers/nfc/port100.c index 5ae61d7ebcfe..30a4e09875d3 100644 --- a/drivers/nfc/port100.c +++ b/drivers/nfc/port100.c @@ -636,6 +636,13 @@ static void port100_recv_response(struct urb *urb) in_frame = dev->in_urb->transfer_buffer; + if (urb->actual_length < PORT100_FRAME_HEADER_LEN || + urb->actual_length < port100_rx_frame_size(in_frame)) { + nfc_err(&dev->interface->dev, "Received a truncated frame\n"); + cmd->status = -EIO; + goto sched_wq; + } + if (!port100_rx_frame_is_valid(in_frame)) { nfc_err(&dev->interface->dev, "Received an invalid frame\n"); cmd->status = -EIO; diff --git a/drivers/nfc/st21nfca/core.c b/drivers/nfc/st21nfca/core.c index fd39a05c9622..b5c1ca3acfbe 100644 --- a/drivers/nfc/st21nfca/core.c +++ b/drivers/nfc/st21nfca/core.c @@ -351,10 +351,10 @@ static int st21nfca_hci_start_poll(struct nfc_hci_dev *hdev, if (r < 0) return r; } else { - hdev->gb = nfc_get_local_general_bytes(hdev->ndev, - &hdev->gb_len); - - if (hdev->gb == NULL || hdev->gb_len == 0) { + nfc_get_local_general_bytes(hdev->ndev, hdev->gb, + sizeof(hdev->gb), + &hdev->gb_len); + if (hdev->gb_len == 0) { im_protocols &= ~NFC_PROTO_NFC_DEP_MASK; tm_protocols &= ~NFC_PROTO_NFC_DEP_MASK; } @@ -577,9 +577,7 @@ static int st21nfca_get_iso15693_inventory(struct nfc_hci_dev *hdev, if (r < 0) goto exit; - skb_pull(inventory_skb, 2); - - if (inventory_skb->len == 0 || + if (!skb_pull(inventory_skb, 2) || inventory_skb->len < 2 || inventory_skb->len > NFC_ISO15693_UID_MAXSIZE) { r = -EPROTO; goto exit; diff --git a/drivers/nfc/st21nfca/i2c.c b/drivers/nfc/st21nfca/i2c.c index aa5f4922b6b0..11ba4fb49828 100644 --- a/drivers/nfc/st21nfca/i2c.c +++ b/drivers/nfc/st21nfca/i2c.c @@ -289,27 +289,36 @@ static int check_crc(u8 *buf, int buflen) */ static int st21nfca_hci_i2c_repack(struct sk_buff *skb) { - int i, j, r, size; + int read, write, r, size; - if (skb->len < 1 || (skb->len > 1 && skb->data[1] != 0)) + if (skb->len < ST21NFCA_FRAME_HEADROOM || + !IS_START_OF_FRAME(skb->data)) return -EBADMSG; size = get_frame_size(skb->data, skb->len); if (size > 0) { + if (size < ST21NFCA_FRAME_HEADROOM + 2) + return -EBADMSG; + skb_trim(skb, size); /* remove ST21NFCA byte stuffing for upper layer */ - for (i = 1, j = 0; i < skb->len; i++) { - if (skb->data[i + j] == + for (read = 1, write = 1; read < skb->len;) { + if (skb->data[read] == (u8) ST21NFCA_ESCAPE_BYTE_STUFFING) { - skb->data[i] = skb->data[i + j + 1] - | ST21NFCA_BYTE_STUFFING_MASK; - i++; - j++; + if (read + 1 == skb->len) + return -EBADMSG; + + skb->data[write++] = skb->data[read + 1] + | ST21NFCA_BYTE_STUFFING_MASK; + read += 2; + } else { + skb->data[write++] = skb->data[read++]; } - skb->data[i] = skb->data[i + j]; } /* remove byte stuffing useless byte */ - skb_trim(skb, i - j); + skb_trim(skb, write); + if (skb->len < ST21NFCA_FRAME_HEADROOM + 2) + return -EBADMSG; /* remove ST21NFCA_SOF_EOF from head */ skb_pull(skb, 1); diff --git a/drivers/nfc/trf7970a.c b/drivers/nfc/trf7970a.c index f22e091019de..3bd8c4636055 100644 --- a/drivers/nfc/trf7970a.c +++ b/drivers/nfc/trf7970a.c @@ -1997,8 +1997,10 @@ static int trf7970a_startup(struct trf7970a *trf) return ret; ret = trf7970a_update_rx_gain_reduction(trf); - if (ret) + if (ret) { + trf7970a_power_down(trf); return ret; + } pm_runtime_set_active(trf->dev); pm_runtime_enable(trf->dev); diff --git a/drivers/nfc/virtual_ncidev.c b/drivers/nfc/virtual_ncidev.c index 8eeb447ac96e..e51c647b27eb 100644 --- a/drivers/nfc/virtual_ncidev.c +++ b/drivers/nfc/virtual_ncidev.c @@ -195,7 +195,8 @@ static const struct file_operations virtual_ncidev_fops = { .write = virtual_ncidev_write, .open = virtual_ncidev_open, .release = virtual_ncidev_close, - .unlocked_ioctl = virtual_ncidev_ioctl + .unlocked_ioctl = virtual_ncidev_ioctl, + .compat_ioctl = compat_ptr_ioctl, }; static struct miscdevice miscdev = { diff --git a/drivers/pci/of_property.c b/drivers/pci/of_property.c index 75a358f73e69..1500740cc55d 100644 --- a/drivers/pci/of_property.c +++ b/drivers/pci/of_property.c @@ -95,9 +95,13 @@ static int of_pci_prop_bus_range(struct pci_dev *pdev, struct of_changeset *ocs, struct device_node *np) { - u32 bus_range[] = { pdev->subordinate->busn_res.start, - pdev->subordinate->busn_res.end }; + u32 bus_range[2]; + if (!pdev->subordinate) + return 0; + + bus_range[0] = pdev->subordinate->busn_res.start; + bus_range[1] = pdev->subordinate->busn_res.end; return of_changeset_add_prop_u32_array(ocs, np, "bus-range", bus_range, ARRAY_SIZE(bus_range)); } @@ -220,6 +224,9 @@ static int of_pci_prop_intr_map(struct pci_dev *pdev, struct of_changeset *ocs, int ret; u8 pin; + if (!pdev->subordinate) + return 0; + pnode = pci_device_to_OF_node(pdev->bus->self); if (!pnode) pnode = pci_bus_to_OF_node(pdev->bus); diff --git a/drivers/pci/setup-bus.c b/drivers/pci/setup-bus.c index c0a949f2c995..9e13489cac71 100644 --- a/drivers/pci/setup-bus.c +++ b/drivers/pci/setup-bus.c @@ -2380,6 +2380,7 @@ int pci_do_resource_release_and_resize(struct pci_dev *pdev, int resno, int size struct resource *res = pci_resource_n(pdev, resno); struct pci_dev_resource *dev_res; struct pci_bus *bus = pdev->bus; + struct pci_dev *bridge = pci_upstream_bridge(pdev); struct resource *b_win, *r; LIST_HEAD(saved); unsigned int i; @@ -2397,6 +2398,8 @@ int pci_do_resource_release_and_resize(struct pci_dev *pdev, int resno, int size if (ret) return ret; + down_read(&pci_bus_sem); + pci_dev_for_each_resource(pdev, r, i) { if (i >= PCI_BRIDGE_RESOURCES) break; @@ -2415,13 +2418,21 @@ int pci_do_resource_release_and_resize(struct pci_dev *pdev, int resno, int size pci_resize_resource_set_size(pdev, resno, size); - if (!bus->self) - goto out; + if (bridge) { + ret = pbus_reassign_bridge_resources(bus, res, &saved); + if (ret) + goto restore; + } else { + /* No bridge window to adjust; let the core reassign the bus. */ + pci_bus_assign_resources(bus); - down_read(&pci_bus_sem); - ret = pbus_reassign_bridge_resources(bus, res, &saved); - if (ret) - goto restore; + list_for_each_entry(dev_res, &saved, list) { + if (!resource_assigned(dev_res->res)) { + ret = -ENOSPC; + goto restore; + } + } + } out: up_read(&pci_bus_sem); diff --git a/drivers/perf/arm_brbe.c b/drivers/perf/arm_brbe.c index ba554e0c846c..254be4da8ae2 100644 --- a/drivers/perf/arm_brbe.c +++ b/drivers/perf/arm_brbe.c @@ -604,7 +604,7 @@ static bool perf_entry_from_brbe_regset(int index, struct perf_branch_entry *ent return false; brbinf = bregs.brbinf; - perf_clear_branch_entry_bitfields(entry); + *entry = (struct perf_branch_entry){ }; if (brbe_record_is_complete(brbinf)) { entry->from = bregs.brbsrc; entry->to = bregs.brbtgt; diff --git a/drivers/pinctrl/meson/pinctrl-meson-s4.c b/drivers/pinctrl/meson/pinctrl-meson-s4.c index 872948699e9f..365dafe457a9 100644 --- a/drivers/pinctrl/meson/pinctrl-meson-s4.c +++ b/drivers/pinctrl/meson/pinctrl-meson-s4.c @@ -854,7 +854,7 @@ static const char * const i2c1_groups[] = { static const char * const i2c2_groups[] = { "i2c2_sda_d", "i2c2_scl_d", "i2c2_sda_h8", "i2c2_scl_h9", - "i2c2_sda_h0", "i2c2_scl_h1l," + "i2c2_sda_h0", "i2c2_scl_h1", }; static const char * const i2c3_groups[] = { diff --git a/drivers/pinctrl/microchip/pinctrl-mpfs-mssio.c b/drivers/pinctrl/microchip/pinctrl-mpfs-mssio.c index ea1026a0d22c..92d38e5336de 100644 --- a/drivers/pinctrl/microchip/pinctrl-mpfs-mssio.c +++ b/drivers/pinctrl/microchip/pinctrl-mpfs-mssio.c @@ -86,7 +86,7 @@ static struct mpfs_pinctrl_bank_voltage mpfs_pinctrl_bank_voltages[8] = { { .uv = 1800000, .val = 4 }, { .uv = 2500000, .val = 6 }, { .uv = 3300000, .val = 8 }, - { .uv = 0, .val = 0x3f }, // pin unused + { .uv = 0, .val = 0xf }, // pin unused }; static int mpfs_pinctrl_get_drive_strength_ma(u32 drive_strength) @@ -156,10 +156,10 @@ static void mpfs_pinctrl_set_bank_voltage(struct mpfs_pinctrl *pctrl, unsigned i u32 val = FIELD_PREP(MPFS_PINCTRL_BANK_VOLTAGE_MASK, bank_voltage); if (pin < MPFS_PINCTRL_BANK2_START) - regmap_assign_bits(pctrl->sysreg_regmap, MPFS_PINCTRL_MSSIO_BANK4_CFG_CR, + regmap_update_bits(pctrl->sysreg_regmap, MPFS_PINCTRL_MSSIO_BANK4_CFG_CR, MPFS_PINCTRL_BANK_VOLTAGE_MASK, val); else - regmap_assign_bits(pctrl->sysreg_regmap, MPFS_PINCTRL_MSSIO_BANK2_CFG_CR, + regmap_update_bits(pctrl->sysreg_regmap, MPFS_PINCTRL_MSSIO_BANK2_CFG_CR, MPFS_PINCTRL_BANK_VOLTAGE_MASK, val); } diff --git a/drivers/pinctrl/pinctrl-generic.c b/drivers/pinctrl/pinctrl-generic.c index fd6bdb74028a..4277c8748513 100644 --- a/drivers/pinctrl/pinctrl-generic.c +++ b/drivers/pinctrl/pinctrl-generic.c @@ -3,8 +3,10 @@ #define pr_fmt(fmt) "generic pinconfig core: " fmt #include +#include #include #include +#include #include #include @@ -196,6 +198,8 @@ static int pinctrl_generic_dt_node_to_map(struct pinctrl_dev *pctldev, int ngroups = 0; int ret; + guard(mutex)(&pctldev->mutex); + *maps = NULL; *num_maps = 0; diff --git a/drivers/pinctrl/pinctrl-single.c b/drivers/pinctrl/pinctrl-single.c index 4d5f85b7e6bb..e0eb7240a985 100644 --- a/drivers/pinctrl/pinctrl-single.c +++ b/drivers/pinctrl/pinctrl-single.c @@ -1628,7 +1628,8 @@ static int pcs_irq_init_chained_handler(struct pcs_device *pcs, &pcs_irqdomain_ops, pcs_soc); if (!pcs->domain) { - irq_set_chained_handler(pcs_soc->irq, NULL); + pcs_irq_free(pcs); + pcs_soc->irq = -1; return -EINVAL; } diff --git a/drivers/pinctrl/qcom/pinctrl-ipq5210.c b/drivers/pinctrl/qcom/pinctrl-ipq5210.c index 827a4ad07a6b..d5b4eb3e7343 100644 --- a/drivers/pinctrl/qcom/pinctrl-ipq5210.c +++ b/drivers/pinctrl/qcom/pinctrl-ipq5210.c @@ -867,6 +867,7 @@ static const struct of_device_id ipq5210_tlmm_of_match[] = { { .compatible = "qcom,ipq5210-tlmm", }, { }, }; +MODULE_DEVICE_TABLE(of, ipq5210_tlmm_of_match); static int ipq5210_tlmm_probe(struct platform_device *pdev) { diff --git a/drivers/pinctrl/qcom/pinctrl-nord.c b/drivers/pinctrl/qcom/pinctrl-nord.c index 7c21306e77ff..7f37f8e819ba 100644 --- a/drivers/pinctrl/qcom/pinctrl-nord.c +++ b/drivers/pinctrl/qcom/pinctrl-nord.c @@ -570,8 +570,10 @@ enum nord_functions { msm_mux_qup0_se5, msm_mux_qup1_se0, msm_mux_qup1_se1, - msm_mux_qup1_se2, - msm_mux_qup1_se3, + msm_mux_qup1_se2_01, + msm_mux_qup1_se2_23, + msm_mux_qup1_se3_01, + msm_mux_qup1_se3_23, msm_mux_qup1_se4, msm_mux_qup1_se5, msm_mux_qup1_se6, @@ -1152,11 +1154,19 @@ static const char *const qup1_se1_groups[] = { "gpio123", "gpio124", "gpio125", "gpio126", }; -static const char *const qup1_se2_groups[] = { - "gpio127", "gpio128", "gpio129", "gpio130", +static const char *const qup1_se2_01_groups[] = { + "gpio127", "gpio128", }; -static const char *const qup1_se3_groups[] = { +static const char *const qup1_se2_23_groups[] = { + "gpio127", "gpio128", +}; + +static const char *const qup1_se3_01_groups[] = { + "gpio129", "gpio130", +}; + +static const char *const qup1_se3_23_groups[] = { "gpio129", "gpio130", }; @@ -1428,8 +1438,10 @@ static const struct pinfunction nord_functions[] = { MSM_PIN_FUNCTION(qup0_se5), MSM_PIN_FUNCTION(qup1_se0), MSM_PIN_FUNCTION(qup1_se1), - MSM_PIN_FUNCTION(qup1_se2), - MSM_PIN_FUNCTION(qup1_se3), + MSM_PIN_FUNCTION(qup1_se2_01), + MSM_PIN_FUNCTION(qup1_se2_23), + MSM_PIN_FUNCTION(qup1_se3_01), + MSM_PIN_FUNCTION(qup1_se3_23), MSM_PIN_FUNCTION(qup1_se4), MSM_PIN_FUNCTION(qup1_se5), MSM_PIN_FUNCTION(qup1_se6), @@ -1633,13 +1645,13 @@ static const struct msm_pingroup nord_groups[] = { _, _, _, _, _, _, _), [126] = PINGROUP(126, qup1_se1, qup1_se0, ccu_i2c_scl, mdp1_vsync_out, _, atest_usb20, ddr_pxi, _, _, _, _), - [127] = PINGROUP(127, qup1_se2, qup1_se2, _, atest_usb21, ddr_pxi, + [127] = PINGROUP(127, qup1_se2_23, qup1_se2_01, _, atest_usb21, ddr_pxi, _, _, _, _, _, _), - [128] = PINGROUP(128, qup1_se2, qup1_se2, _, atest_usb20, ddr_pxi, + [128] = PINGROUP(128, qup1_se2_23, qup1_se2_01, _, atest_usb20, ddr_pxi, _, _, _, _, _, _), - [129] = PINGROUP(129, qup1_se3, qup1_se3, ccu_i2c_sda, mdp1_vsync_out, + [129] = PINGROUP(129, qup1_se3_23, qup1_se3_01, ccu_i2c_sda, mdp1_vsync_out, _, atest_usb21, ddr_pxi, _, _, _, _), - [130] = PINGROUP(130, qup1_se3, qup1_se3, ccu_i2c_scl, mdp1_vsync_out, + [130] = PINGROUP(130, qup1_se3_23, qup1_se3_01, ccu_i2c_scl, mdp1_vsync_out, _, atest_usb20, ddr_pxi, _, _, _, _), [131] = PINGROUP(131, qup1_se4, qup1_se6, ccu_i2c_sda, mdp1_vsync_out, _, atest_usb21, ddr_pxi, _, _, _, _), diff --git a/drivers/pinctrl/sunxi/pinctrl-sun55i-a523-r.c b/drivers/pinctrl/sunxi/pinctrl-sun55i-a523-r.c index 462aa1c4a5fa..e27e4945def2 100644 --- a/drivers/pinctrl/sunxi/pinctrl-sun55i-a523-r.c +++ b/drivers/pinctrl/sunxi/pinctrl-sun55i-a523-r.c @@ -26,7 +26,7 @@ static const u8 a523_r_irq_bank_muxes[SUNXI_PINCTRL_MAX_BANKS] = static struct sunxi_pinctrl_desc a523_r_pinctrl_data = { .irq_banks = ARRAY_SIZE(a523_r_irq_bank_map), .irq_bank_map = a523_r_irq_bank_map, - .io_bias_cfg_variant = BIAS_VOLTAGE_PIO_POW_MODE_SEL, + .io_bias_cfg_variant = BIAS_VOLTAGE_PIO_POW_MODE_CTL_INV, .pin_base = PL_BASE, }; diff --git a/drivers/pinctrl/sunxi/pinctrl-sun55i-a523.c b/drivers/pinctrl/sunxi/pinctrl-sun55i-a523.c index b6f78f1f30ac..88d8acd5bc24 100644 --- a/drivers/pinctrl/sunxi/pinctrl-sun55i-a523.c +++ b/drivers/pinctrl/sunxi/pinctrl-sun55i-a523.c @@ -26,7 +26,7 @@ static const u8 a523_irq_bank_muxes[SUNXI_PINCTRL_MAX_BANKS] = static struct sunxi_pinctrl_desc a523_pinctrl_data = { .irq_banks = ARRAY_SIZE(a523_irq_bank_map), .irq_bank_map = a523_irq_bank_map, - .io_bias_cfg_variant = BIAS_VOLTAGE_PIO_POW_MODE_SEL, + .io_bias_cfg_variant = BIAS_VOLTAGE_PIO_POW_MODE_CTL_INV, }; static int a523_pinctrl_probe(struct platform_device *pdev) diff --git a/drivers/pinctrl/sunxi/pinctrl-sunxi.c b/drivers/pinctrl/sunxi/pinctrl-sunxi.c index 25489beeb312..31cd142ce0f7 100644 --- a/drivers/pinctrl/sunxi/pinctrl-sunxi.c +++ b/drivers/pinctrl/sunxi/pinctrl-sunxi.c @@ -728,6 +728,7 @@ static int sunxi_pinctrl_set_io_bias_cfg(struct sunxi_pinctrl *pctl, { unsigned short bank; unsigned long flags; + bool inverted = false; u32 val, reg; int uV; @@ -766,6 +767,9 @@ static int sunxi_pinctrl_set_io_bias_cfg(struct sunxi_pinctrl *pctl, reg &= ~IO_BIAS_MASK; writel(reg | val, pctl->membase + sunxi_grp_config_reg(pin)); return 0; + case BIAS_VOLTAGE_PIO_POW_MODE_CTL_INV: + inverted = true; + fallthrough; case BIAS_VOLTAGE_PIO_POW_MODE_CTL: val = uV > 1800000 && uV <= 2500000 ? BIT(bank) : 0; @@ -780,6 +784,8 @@ static int sunxi_pinctrl_set_io_bias_cfg(struct sunxi_pinctrl *pctl, fallthrough; case BIAS_VOLTAGE_PIO_POW_MODE_SEL: val = uV <= 1800000 ? 1 : 0; + if (inverted) + val = !val; raw_spin_lock_irqsave(&pctl->lock, flags); reg = readl(pctl->membase + pctl->pow_mod_sel_offset); @@ -837,6 +843,21 @@ static void sunxi_pmx_set(struct pinctrl_dev *pctldev, writel((readl(pctl->membase + reg) & ~mask) | config << shift, pctl->membase + reg); + /* + * A pin muxed to gpio_out directly through a pinmux node bypasses + * sunxi_pinctrl_gpio_set() and drives whatever its output latch + * holds. Now that the pin is in output mode the data register + * reads back the latch, so refresh the shadow to keep such pins + * driving their pre-existing level. + */ + if (config == SUN4I_FUNC_OUTPUT) { + u32 *shadow = &pctl->dat_shadow[pin / PINS_PER_BANK]; + + sunxi_data_reg(pctl, pin, ®, &shift, &mask); + *shadow = (*shadow & ~mask) | + (readl(pctl->membase + reg) & mask); + } + raw_spin_unlock_irqrestore(&pctl->lock, flags); } @@ -1017,21 +1038,29 @@ static int sunxi_pinctrl_gpio_set(struct gpio_chip *chip, unsigned int offset, int value) { struct sunxi_pinctrl *pctl = gpiochip_get_data(chip); - u32 reg, shift, mask, val; + u32 *shadow = &pctl->dat_shadow[offset / PINS_PER_BANK]; + u32 reg, shift, mask; unsigned long flags; sunxi_data_reg(pctl, offset, ®, &shift, &mask); raw_spin_lock_irqsave(&pctl->lock, flags); - val = readl(pctl->membase + reg); - + /* + * Reading the data register returns the pin level, not the output + * latch, for pins muxed as inputs. A read-modify-write based on + * the register would therefore corrupt the latches of input-muxed + * pins in the same bank (e.g. an emulated open-drain I2C line + * released high), making them drive the wrong level once switched + * to output. Base the read-modify-write on a shadow copy of the + * latches instead. + */ if (value) - val |= mask; + *shadow |= mask; else - val &= ~mask; + *shadow &= ~mask; - writel(val, pctl->membase + reg); + writel(*shadow, pctl->membase + reg); raw_spin_unlock_irqrestore(&pctl->lock, flags); @@ -1572,7 +1601,7 @@ int sunxi_pinctrl_init_with_flags(struct platform_device *pdev, struct pinctrl_pin_desc *pins; struct sunxi_pinctrl *pctl; struct pinmux_ops *pmxops; - int i, ret, last_pin, pin_idx; + int i, ret, last_pin, pin_idx, nbanks; struct clk *clk; pctl = devm_kzalloc(&pdev->dev, sizeof(*pctl), GFP_KERNEL); @@ -1610,6 +1639,37 @@ int sunxi_pinctrl_init_with_flags(struct platform_device *pdev, if (!pctl->irq_array) return -ENOMEM; + /* + * The bus clock has to be enabled before the pinctrl device + * registers, as the pin hogs claimed from there access registers. + */ + ret = of_clk_get_parent_count(node); + clk = devm_clk_get_enabled(&pdev->dev, ret == 1 ? NULL : "apb"); + if (IS_ERR(clk)) + return PTR_ERR(clk); + + /* + * Seed the output latch shadow from the hardware so pins the + * bootloader left in output mode keep their state; see + * sunxi_pinctrl_gpio_set() for why a shadow is needed. This must + * happen before the pinctrl device registers, as pin hogs can mux + * pins to gpio_out and thereby update the shadow. + */ + last_pin = pctl->desc->pins[pctl->desc->npins - 1].pin.number; + nbanks = DIV_ROUND_UP(last_pin + 1 - pctl->desc->pin_base, + PINS_PER_BANK); + pctl->dat_shadow = devm_kcalloc(&pdev->dev, nbanks, + sizeof(*pctl->dat_shadow), GFP_KERNEL); + if (!pctl->dat_shadow) + return -ENOMEM; + + for (i = 0; i < nbanks; i++) { + u32 reg, shift, mask; + + sunxi_data_reg(pctl, i * PINS_PER_BANK, ®, &shift, &mask); + pctl->dat_shadow[i] = readl(pctl->membase + reg); + } + ret = sunxi_pinctrl_build_state(pdev); if (ret) { dev_err(&pdev->dev, "dt probe failed: %d\n", ret); @@ -1665,7 +1725,6 @@ int sunxi_pinctrl_init_with_flags(struct platform_device *pdev, if (!pctl->chip) return -ENOMEM; - last_pin = pctl->desc->pins[pctl->desc->npins - 1].pin.number; pctl->chip->owner = THIS_MODULE; pctl->chip->request = gpiochip_generic_request; pctl->chip->free = gpiochip_generic_free; @@ -1699,13 +1758,6 @@ int sunxi_pinctrl_init_with_flags(struct platform_device *pdev, goto gpiochip_error; } - ret = of_clk_get_parent_count(node); - clk = devm_clk_get_enabled(&pdev->dev, ret == 1 ? NULL : "apb"); - if (IS_ERR(clk)) { - ret = PTR_ERR(clk); - goto gpiochip_error; - } - pctl->irq = devm_kcalloc(&pdev->dev, pctl->desc->irq_banks, sizeof(*pctl->irq), diff --git a/drivers/pinctrl/sunxi/pinctrl-sunxi.h b/drivers/pinctrl/sunxi/pinctrl-sunxi.h index 0daf7600e2fb..8bd00c6ff628 100644 --- a/drivers/pinctrl/sunxi/pinctrl-sunxi.h +++ b/drivers/pinctrl/sunxi/pinctrl-sunxi.h @@ -85,6 +85,7 @@ #define IO_BIAS_MASK GENMASK(3, 0) #define SUN4I_FUNC_INPUT 0 +#define SUN4I_FUNC_OUTPUT 1 #define SUN4I_FUNC_IRQ 6 #define SUN4I_FUNC_DISABLED_OLD 7 #define SUN4I_FUNC_DISABLED_NEW 15 @@ -116,8 +117,10 @@ enum sunxi_desc_bias_voltage { * Bias voltage is set through PIO_POW_MOD_SEL_REG * and PIO_POW_MOD_CTL_REG register, as seen on * A100 and D1 SoC, for example. + * Some SoCs invert the encoding for 1.8V vs. 3.3V. */ BIAS_VOLTAGE_PIO_POW_MODE_CTL, + BIAS_VOLTAGE_PIO_POW_MODE_CTL_INV, }; struct sunxi_desc_function { @@ -175,6 +178,12 @@ struct sunxi_pinctrl { int *irq; unsigned *irq_array; raw_spinlock_t lock; + /* + * Output latch shadow, one word per bank. Seeded lockless at + * probe before the pinctrl device registers, protected by @lock + * afterwards. + */ + u32 *dat_shadow; struct pinctrl_dev *pctl_dev; unsigned long flags; u32 bank_mem_size; diff --git a/drivers/pinctrl/tegra/pinctrl-tegra238.c b/drivers/pinctrl/tegra/pinctrl-tegra238.c index ec482365f14f..40aba285944e 100644 --- a/drivers/pinctrl/tegra/pinctrl-tegra238.c +++ b/drivers/pinctrl/tegra/pinctrl-tegra238.c @@ -1744,57 +1744,57 @@ static const char * const tegra238_functions[] = { #define drive_sdmmc1_dat0_pu2 DRV_PINGROUP_ENTRY_Y(0x8034, 28, 2, 30, 2, -1, -1, -1, -1, 0) #define drive_ufs0_rst_n_pv1 DRV_PINGROUP_ENTRY_Y(0x11004, 12, 5, 24, 5, -1, -1, -1, -1, 0) #define drive_ufs0_ref_clk_pv0 DRV_PINGROUP_ENTRY_Y(0x1100c, 12, 5, 24, 5, -1, -1, -1, -1, 0) -#define drive_batt_oc_paa4 DRV_PINGROUP_ENTRY_Y(0x1024, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_bootv_ctl_n_paa0 DRV_PINGROUP_ENTRY_Y(0x102c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_vcomp_alert_paa2 DRV_PINGROUP_ENTRY_Y(0x105c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_hdmi_cec_pbb0 DRV_PINGROUP_ENTRY_Y(0x1064, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_touch_clk_pdd3 DRV_PINGROUP_ENTRY_Y(0x106c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart3_rx_pcc6 DRV_PINGROUP_ENTRY_Y(0x1074, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart3_tx_pcc5 DRV_PINGROUP_ENTRY_Y(0x107c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_gen8_i2c_sda_pdd2 DRV_PINGROUP_ENTRY_Y(0x1084, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_gen8_i2c_scl_pdd1 DRV_PINGROUP_ENTRY_Y(0x108c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_spi2_mosi_pcc2 DRV_PINGROUP_ENTRY_Y(0x1094, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_gen2_i2c_scl_pcc7 DRV_PINGROUP_ENTRY_Y(0x109c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_spi2_cs0_pcc3 DRV_PINGROUP_ENTRY_Y(0x10a4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_gen2_i2c_sda_pdd0 DRV_PINGROUP_ENTRY_Y(0x10ac, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_spi2_sck_pcc0 DRV_PINGROUP_ENTRY_Y(0x10b4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_spi2_miso_pcc1 DRV_PINGROUP_ENTRY_Y(0x10bc, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio49_pee2 DRV_PINGROUP_ENTRY_Y(0x10c4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio50_pee4 DRV_PINGROUP_ENTRY_Y(0x10cc, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio82_pee3 DRV_PINGROUP_ENTRY_Y(0x10d4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio71_pff2 DRV_PINGROUP_ENTRY_Y(0x10dc, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio76_pff7 DRV_PINGROUP_ENTRY_Y(0x10e4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio74_pff5 DRV_PINGROUP_ENTRY_Y(0x10ec, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio00_paa1 DRV_PINGROUP_ENTRY_Y(0x10f4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio19_pdd6 DRV_PINGROUP_ENTRY_Y(0x10fc, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio86_phh3 DRV_PINGROUP_ENTRY_Y(0x1104, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio72_pff3 DRV_PINGROUP_ENTRY_Y(0x110c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio77_pgg0 DRV_PINGROUP_ENTRY_Y(0x1114, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio80_pff6 DRV_PINGROUP_ENTRY_Y(0x111c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio84_pgg1 DRV_PINGROUP_ENTRY_Y(0x1124, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio83_pee5 DRV_PINGROUP_ENTRY_Y(0x112c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio73_pff4 DRV_PINGROUP_ENTRY_Y(0x1134, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio70_pff1 DRV_PINGROUP_ENTRY_Y(0x113c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio04_paa5 DRV_PINGROUP_ENTRY_Y(0x1144, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio85_pgg6 DRV_PINGROUP_ENTRY_Y(0x114c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio69_pff0 DRV_PINGROUP_ENTRY_Y(0x1154, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio25_paa6 DRV_PINGROUP_ENTRY_Y(0x115c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_soc_gpio26_paa7 DRV_PINGROUP_ENTRY_Y(0x1164, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart5_tx_pgg7 DRV_PINGROUP_ENTRY_Y(0x116c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart5_rx_phh0 DRV_PINGROUP_ENTRY_Y(0x1174, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart2_tx_pgg2 DRV_PINGROUP_ENTRY_Y(0x117c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart2_rx_pgg3 DRV_PINGROUP_ENTRY_Y(0x1184, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart2_cts_pgg5 DRV_PINGROUP_ENTRY_Y(0x118c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart2_rts_pgg4 DRV_PINGROUP_ENTRY_Y(0x1194, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart5_cts_phh2 DRV_PINGROUP_ENTRY_Y(0x119c, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_uart5_rts_phh1 DRV_PINGROUP_ENTRY_Y(0x11a4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_pwm7_pee1 DRV_PINGROUP_ENTRY_Y(0x11ac, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_pwm2_pdd7 DRV_PINGROUP_ENTRY_Y(0x11b4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_pwm3_pee0 DRV_PINGROUP_ENTRY_Y(0x11bc, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_pwm1_paa3 DRV_PINGROUP_ENTRY_Y(0x11c4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_spi2_cs1_pcc4 DRV_PINGROUP_ENTRY_Y(0x11cc, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_dmic1_clk_pdd4 DRV_PINGROUP_ENTRY_Y(0x11d4, 12, 5, 20, 5, -1, -1, -1, -1, 1) -#define drive_dmic1_dat_pdd5 DRV_PINGROUP_ENTRY_Y(0x11dc, 12, 5, 20, 5, -1, -1, -1, -1, 1) +#define drive_batt_oc_paa4 DRV_PINGROUP_ENTRY_Y(0x1024, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_bootv_ctl_n_paa0 DRV_PINGROUP_ENTRY_Y(0x102c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_vcomp_alert_paa2 DRV_PINGROUP_ENTRY_Y(0x105c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_hdmi_cec_pbb0 DRV_PINGROUP_ENTRY_Y(0x1064, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_touch_clk_pdd3 DRV_PINGROUP_ENTRY_Y(0x106c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart3_rx_pcc6 DRV_PINGROUP_ENTRY_Y(0x1074, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart3_tx_pcc5 DRV_PINGROUP_ENTRY_Y(0x107c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_gen8_i2c_sda_pdd2 DRV_PINGROUP_ENTRY_Y(0x1084, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_gen8_i2c_scl_pdd1 DRV_PINGROUP_ENTRY_Y(0x108c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_spi2_mosi_pcc2 DRV_PINGROUP_ENTRY_Y(0x1094, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_gen2_i2c_scl_pcc7 DRV_PINGROUP_ENTRY_Y(0x109c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_spi2_cs0_pcc3 DRV_PINGROUP_ENTRY_Y(0x10a4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_gen2_i2c_sda_pdd0 DRV_PINGROUP_ENTRY_Y(0x10ac, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_spi2_sck_pcc0 DRV_PINGROUP_ENTRY_Y(0x10b4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_spi2_miso_pcc1 DRV_PINGROUP_ENTRY_Y(0x10bc, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio49_pee2 DRV_PINGROUP_ENTRY_Y(0x10c4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio50_pee4 DRV_PINGROUP_ENTRY_Y(0x10cc, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio82_pee3 DRV_PINGROUP_ENTRY_Y(0x10d4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio71_pff2 DRV_PINGROUP_ENTRY_Y(0x10dc, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio76_pff7 DRV_PINGROUP_ENTRY_Y(0x10e4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio74_pff5 DRV_PINGROUP_ENTRY_Y(0x10ec, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio00_paa1 DRV_PINGROUP_ENTRY_Y(0x10f4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio19_pdd6 DRV_PINGROUP_ENTRY_Y(0x10fc, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio86_phh3 DRV_PINGROUP_ENTRY_Y(0x1104, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio72_pff3 DRV_PINGROUP_ENTRY_Y(0x110c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio77_pgg0 DRV_PINGROUP_ENTRY_Y(0x1114, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio80_pff6 DRV_PINGROUP_ENTRY_Y(0x111c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio84_pgg1 DRV_PINGROUP_ENTRY_Y(0x1124, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio83_pee5 DRV_PINGROUP_ENTRY_Y(0x112c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio73_pff4 DRV_PINGROUP_ENTRY_Y(0x1134, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio70_pff1 DRV_PINGROUP_ENTRY_Y(0x113c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio04_paa5 DRV_PINGROUP_ENTRY_Y(0x1144, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio85_pgg6 DRV_PINGROUP_ENTRY_Y(0x114c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio69_pff0 DRV_PINGROUP_ENTRY_Y(0x1154, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio25_paa6 DRV_PINGROUP_ENTRY_Y(0x115c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_soc_gpio26_paa7 DRV_PINGROUP_ENTRY_Y(0x1164, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart5_tx_pgg7 DRV_PINGROUP_ENTRY_Y(0x116c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart5_rx_phh0 DRV_PINGROUP_ENTRY_Y(0x1174, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart2_tx_pgg2 DRV_PINGROUP_ENTRY_Y(0x117c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart2_rx_pgg3 DRV_PINGROUP_ENTRY_Y(0x1184, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart2_cts_pgg5 DRV_PINGROUP_ENTRY_Y(0x118c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart2_rts_pgg4 DRV_PINGROUP_ENTRY_Y(0x1194, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart5_cts_phh2 DRV_PINGROUP_ENTRY_Y(0x119c, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_uart5_rts_phh1 DRV_PINGROUP_ENTRY_Y(0x11a4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_pwm7_pee1 DRV_PINGROUP_ENTRY_Y(0x11ac, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_pwm2_pdd7 DRV_PINGROUP_ENTRY_Y(0x11b4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_pwm3_pee0 DRV_PINGROUP_ENTRY_Y(0x11bc, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_pwm1_paa3 DRV_PINGROUP_ENTRY_Y(0x11c4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_spi2_cs1_pcc4 DRV_PINGROUP_ENTRY_Y(0x11cc, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_dmic1_clk_pdd4 DRV_PINGROUP_ENTRY_Y(0x11d4, 12, 5, 20, 5, -1, -1, -1, -1, 0) +#define drive_dmic1_dat_pdd5 DRV_PINGROUP_ENTRY_Y(0x11dc, 12, 5, 20, 5, -1, -1, -1, -1, 0) #define drive_sdmmc1_comp DRV_PINGROUP_ENTRY_N @@ -1961,57 +1961,57 @@ static const struct tegra_pingroup tegra238_groups[] = { }; static const struct tegra_pingroup tegra238_aon_groups[] = { - PINGROUP(bootv_ctl_n_paa0, RSVD0, RSVD1, RSVD2, RSVD3, 0x1028, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio00_paa1, RSVD0, RSVD1, RSVD2, RSVD3, 0x10f0, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(vcomp_alert_paa2, SOC_THERM_OC1, RSVD1, RSVD2, RSVD3, 0x1058, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(pwm1_paa3, GP_PWM1, RSVD1, RSVD2, RSVD3, 0x11c0, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(batt_oc_paa4, SOC_THERM_OC2, RSVD1, RSVD2, RSVD3, 0x1020, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio04_paa5, RSVD0, RSVD1, RSVD2, RSVD3, 0x1140, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio25_paa6, RSVD0, RSVD1, RSVD2, RSVD3, 0x1158, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio26_paa7, RSVD0, SOC_THERM_OC3, RSVD2, RSVD3, 0x1160, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(hdmi_cec_pbb0, HDMI_CEC, RSVD1, RSVD2, RSVD3, 0x1060, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(spi2_sck_pcc0, SPI2_SCK, RSVD1, RSVD2, RSVD3, 0x10b0, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(spi2_miso_pcc1, SPI2_DIN, RSVD1, RSVD2, RSVD3, 0x10b8, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(spi2_mosi_pcc2, SPI2_DOUT, RSVD1, RSVD2, RSVD3, 0x1090, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(spi2_cs0_pcc3, SPI2_CS0, RSVD1, RSVD2, RSVD3, 0x10a0, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(spi2_cs1_pcc4, SPI2_CS1, RSVD1, RSVD2, RSVD3, 0x11c8, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(uart3_tx_pcc5, UARTC_TXD, RSVD1, RSVD2, RSVD3, 0x1078, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart3_rx_pcc6, UARTC_RXD, RSVD1, RSVD2, RSVD3, 0x1070, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(gen2_i2c_scl_pcc7, I2C2_CLK, RSVD1, RSVD2, RSVD3, 0x1098, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(gen2_i2c_sda_pdd0, I2C2_DAT, RSVD1, RSVD2, RSVD3, 0x10a8, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(gen8_i2c_scl_pdd1, I2C8_CLK, RSVD1, RSVD2, RSVD3, 0x1088, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(gen8_i2c_sda_pdd2, I2C8_DAT, RSVD1, RSVD2, RSVD3, 0x1080, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(touch_clk_pdd3, GP_PWM4, TOUCH_CLK, RSVD2, RSVD3, 0x1068, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(dmic1_clk_pdd4, DMIC1_CLK, RSVD1, DMIC5_CLK, RSVD3, 0x11d0, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(dmic1_dat_pdd5, DMIC1_DAT, RSVD1, DMIC5_DAT, RSVD3, 0x11d8, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio19_pdd6, RSVD0, WDT_RESET_OUTB, RSVD2, RSVD3, 0x10f8, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio49_pee2, RSVD0, RSVD1, RSVD2, RSVD3, 0x10c0, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio50_pee4, RSVD0, RSVD1, RSVD2, RSVD3, 0x10c8, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio82_pee3, RSVD0, RSVD1, RSVD2, RSVD3, 0x10d0, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio71_pff2, PPC_MODE_1, RSVD1, RSVD2, RSVD3, 0x10d8, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio76_pff7, RSVD0, RSVD1, TSC_EDGE_OUT0, TSC_EDGE_OUT0A, 0x10e0, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio74_pff5, PPC_READY, PPC_I2C_DAT, RSVD2, RSVD3, 0x10e8, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio86_phh3, RSVD0, SPI5_CS1, TSC_EDGE_OUT3, TSC_EDGE_OUT0D, 0x1100, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio72_pff3, PPC_MODE_2, RSVD1, RSVD2, RSVD3, 0x1108, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio77_pgg0, RSVD0, RSVD1, TSC_EDGE_OUT1, TSC_EDGE_OUT0B, 0x1110, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio80_pff6, RSVD0, PPC_RST_N, RSVD2, RSVD3, 0x1118, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio84_pgg1, RSVD0, RSVD1, TSC_EDGE_OUT2, TSC_EDGE_OUT0C, 0x1120, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio83_pee5, RSVD0, RSVD1, RSVD2, RSVD3, 0x1128, 1, Y, -1, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio73_pff4, PPC_CC, PPC_I2C_CLK, RSVD2, RSVD3, 0x1130, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio70_pff1, PPC_MODE_0, RSVD1, RSVD2, RSVD3, 0x1138, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio85_pgg6, RSVD0, SPI4_CS1, RSVD2, RSVD3, 0x1148, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(soc_gpio69_pff0, PPC_INT_N, RSVD1, RSVD2, RSVD3, 0x1150, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart5_tx_pgg7, UARTE_TXD, SPI5_SCK, RSVD2, RSVD3, 0x1168, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart5_rx_phh0, UARTE_RXD, SPI5_MISO, RSVD2, RSVD3, 0x1170, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart2_tx_pgg2, UARTB_TXD, SPI4_SCK, RSVD2, RSVD3, 0x1178, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart2_rx_pgg3, UARTB_RXD, SPI4_MISO, RSVD2, RSVD3, 0x1180, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart2_cts_pgg5, UARTB_CTS, SPI4_CS0, RSVD2, RSVD3, 0x1188, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart2_rts_pgg4, UARTB_RTS, SPI4_MOSI, RSVD2, RSVD3, 0x1190, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart5_cts_phh2, UARTE_CTS, SPI5_CS0, RSVD2, RSVD3, 0x1198, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(uart5_rts_phh1, UARTE_RTS, SPI5_MOSI, RSVD2, RSVD3, 0x11a0, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(pwm2_pdd7, GP_PWM2, LED_BLINK, RSVD2, RSVD3, 0x11b0, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(pwm3_pee0, GP_PWM3, RSVD1, RSVD2, RSVD3, 0x11b8, 1, Y, 5, 7, 6, 8, -1, 10, 12), - PINGROUP(pwm7_pee1, GP_PWM7, RSVD1, RSVD2, RSVD3, 0x11a8, 1, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(bootv_ctl_n_paa0, RSVD0, RSVD1, RSVD2, RSVD3, 0x1028, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio00_paa1, RSVD0, RSVD1, RSVD2, RSVD3, 0x10f0, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(vcomp_alert_paa2, SOC_THERM_OC1, RSVD1, RSVD2, RSVD3, 0x1058, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(pwm1_paa3, GP_PWM1, RSVD1, RSVD2, RSVD3, 0x11c0, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(batt_oc_paa4, SOC_THERM_OC2, RSVD1, RSVD2, RSVD3, 0x1020, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio04_paa5, RSVD0, RSVD1, RSVD2, RSVD3, 0x1140, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio25_paa6, RSVD0, RSVD1, RSVD2, RSVD3, 0x1158, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio26_paa7, RSVD0, SOC_THERM_OC3, RSVD2, RSVD3, 0x1160, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(hdmi_cec_pbb0, HDMI_CEC, RSVD1, RSVD2, RSVD3, 0x1060, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(spi2_sck_pcc0, SPI2_SCK, RSVD1, RSVD2, RSVD3, 0x10b0, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(spi2_miso_pcc1, SPI2_DIN, RSVD1, RSVD2, RSVD3, 0x10b8, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(spi2_mosi_pcc2, SPI2_DOUT, RSVD1, RSVD2, RSVD3, 0x1090, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(spi2_cs0_pcc3, SPI2_CS0, RSVD1, RSVD2, RSVD3, 0x10a0, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(spi2_cs1_pcc4, SPI2_CS1, RSVD1, RSVD2, RSVD3, 0x11c8, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(uart3_tx_pcc5, UARTC_TXD, RSVD1, RSVD2, RSVD3, 0x1078, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart3_rx_pcc6, UARTC_RXD, RSVD1, RSVD2, RSVD3, 0x1070, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(gen2_i2c_scl_pcc7, I2C2_CLK, RSVD1, RSVD2, RSVD3, 0x1098, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(gen2_i2c_sda_pdd0, I2C2_DAT, RSVD1, RSVD2, RSVD3, 0x10a8, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(gen8_i2c_scl_pdd1, I2C8_CLK, RSVD1, RSVD2, RSVD3, 0x1088, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(gen8_i2c_sda_pdd2, I2C8_DAT, RSVD1, RSVD2, RSVD3, 0x1080, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(touch_clk_pdd3, GP_PWM4, TOUCH_CLK, RSVD2, RSVD3, 0x1068, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(dmic1_clk_pdd4, DMIC1_CLK, RSVD1, DMIC5_CLK, RSVD3, 0x11d0, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(dmic1_dat_pdd5, DMIC1_DAT, RSVD1, DMIC5_DAT, RSVD3, 0x11d8, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio19_pdd6, RSVD0, WDT_RESET_OUTB, RSVD2, RSVD3, 0x10f8, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio49_pee2, RSVD0, RSVD1, RSVD2, RSVD3, 0x10c0, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio50_pee4, RSVD0, RSVD1, RSVD2, RSVD3, 0x10c8, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio82_pee3, RSVD0, RSVD1, RSVD2, RSVD3, 0x10d0, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio71_pff2, PPC_MODE_1, RSVD1, RSVD2, RSVD3, 0x10d8, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio76_pff7, RSVD0, RSVD1, TSC_EDGE_OUT0, TSC_EDGE_OUT0A, 0x10e0, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio74_pff5, PPC_READY, PPC_I2C_DAT, RSVD2, RSVD3, 0x10e8, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio86_phh3, RSVD0, SPI5_CS1, TSC_EDGE_OUT3, TSC_EDGE_OUT0D, 0x1100, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio72_pff3, PPC_MODE_2, RSVD1, RSVD2, RSVD3, 0x1108, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio77_pgg0, RSVD0, RSVD1, TSC_EDGE_OUT1, TSC_EDGE_OUT0B, 0x1110, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio80_pff6, RSVD0, PPC_RST_N, RSVD2, RSVD3, 0x1118, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio84_pgg1, RSVD0, RSVD1, TSC_EDGE_OUT2, TSC_EDGE_OUT0C, 0x1120, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio83_pee5, RSVD0, RSVD1, RSVD2, RSVD3, 0x1128, 0, Y, -1, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio73_pff4, PPC_CC, PPC_I2C_CLK, RSVD2, RSVD3, 0x1130, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio70_pff1, PPC_MODE_0, RSVD1, RSVD2, RSVD3, 0x1138, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio85_pgg6, RSVD0, SPI4_CS1, RSVD2, RSVD3, 0x1148, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(soc_gpio69_pff0, PPC_INT_N, RSVD1, RSVD2, RSVD3, 0x1150, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart5_tx_pgg7, UARTE_TXD, SPI5_SCK, RSVD2, RSVD3, 0x1168, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart5_rx_phh0, UARTE_RXD, SPI5_MISO, RSVD2, RSVD3, 0x1170, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart2_tx_pgg2, UARTB_TXD, SPI4_SCK, RSVD2, RSVD3, 0x1178, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart2_rx_pgg3, UARTB_RXD, SPI4_MISO, RSVD2, RSVD3, 0x1180, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart2_cts_pgg5, UARTB_CTS, SPI4_CS0, RSVD2, RSVD3, 0x1188, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart2_rts_pgg4, UARTB_RTS, SPI4_MOSI, RSVD2, RSVD3, 0x1190, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart5_cts_phh2, UARTE_CTS, SPI5_CS0, RSVD2, RSVD3, 0x1198, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(uart5_rts_phh1, UARTE_RTS, SPI5_MOSI, RSVD2, RSVD3, 0x11a0, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(pwm2_pdd7, GP_PWM2, LED_BLINK, RSVD2, RSVD3, 0x11b0, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(pwm3_pee0, GP_PWM3, RSVD1, RSVD2, RSVD3, 0x11b8, 0, Y, 5, 7, 6, 8, -1, 10, 12), + PINGROUP(pwm7_pee1, GP_PWM7, RSVD1, RSVD2, RSVD3, 0x11a8, 0, Y, 5, 7, 6, 8, -1, 10, 12), }; static const struct tegra_pinctrl_soc_data tegra238_pinctrl_aon = { diff --git a/drivers/remoteproc/qcom_q6v5_adsp.c b/drivers/remoteproc/qcom_q6v5_adsp.c index c81e6c33c747..9c5dedf0aa60 100644 --- a/drivers/remoteproc/qcom_q6v5_adsp.c +++ b/drivers/remoteproc/qcom_q6v5_adsp.c @@ -104,6 +104,7 @@ struct qcom_adsp { struct completion stop_done; phys_addr_t mem_phys; + unsigned long iova; phys_addr_t mem_reloc; void *mem_region; size_t mem_size; @@ -333,7 +334,7 @@ static void adsp_unmap_carveout(struct rproc *rproc) struct qcom_adsp *adsp = rproc->priv; if (adsp->has_iommu) - iommu_unmap(rproc->domain, adsp->mem_phys, adsp->mem_size); + iommu_unmap(rproc->domain, adsp->iova, adsp->mem_size); } static int adsp_map_carveout(struct rproc *rproc) @@ -341,7 +342,6 @@ static int adsp_map_carveout(struct rproc *rproc) struct qcom_adsp *adsp = rproc->priv; struct of_phandle_args args; long long sid; - unsigned long iova; int ret; if (!adsp->has_iommu) @@ -358,9 +358,9 @@ static int adsp_map_carveout(struct rproc *rproc) of_node_put(args.np); /* Add SID configuration for ADSP Firmware to SMMU */ - iova = adsp->mem_phys | (sid << 32); + adsp->iova = adsp->mem_phys | (sid << 32); - ret = iommu_map(rproc->domain, iova, adsp->mem_phys, + ret = iommu_map(rproc->domain, adsp->iova, adsp->mem_phys, adsp->mem_size, IOMMU_READ | IOMMU_WRITE, GFP_KERNEL); if (ret) { diff --git a/drivers/s390/cio/chp.c b/drivers/s390/cio/chp.c index c890f21a82ce..eaf0527bff6c 100644 --- a/drivers/s390/cio/chp.c +++ b/drivers/s390/cio/chp.c @@ -78,6 +78,9 @@ u8 chp_get_sch_opm(struct subchannel *sch) int opm; int i; + if (!sch->schib.pmcw.dnv) + return 0; + opm = 0; chp_id_init(&chpid); for (i = 0; i < 8; i++) { diff --git a/drivers/s390/cio/cio.c b/drivers/s390/cio/cio.c index 70dc8cc76594..e1c62eb60cca 100644 --- a/drivers/s390/cio/cio.c +++ b/drivers/s390/cio/cio.c @@ -453,7 +453,8 @@ EXPORT_SYMBOL_GPL(cio_commit_config); /** * cio_update_schib - Perform stsch and update schib if subchannel is valid. * @sch: subchannel on which to perform stsch - * Return zero on success, -ENODEV otherwise. + * Return zero on success, -ENODEV if the subchannel is not operational, + * -EACCES if the subchannel has no valid device. */ int cio_update_schib(struct subchannel *sch) { @@ -462,10 +463,12 @@ int cio_update_schib(struct subchannel *sch) if (stsch(sch->schid, &schib)) return -ENODEV; - memcpy(&sch->schib, &schib, sizeof(schib)); - - if (!css_sch_is_valid(&schib)) + if (!css_sch_is_valid(&schib)) { + memset(&sch->schib, 0, sizeof(sch->schib)); return -EACCES; + } + + memcpy(&sch->schib, &schib, sizeof(schib)); return 0; } diff --git a/drivers/s390/cio/cio.h b/drivers/s390/cio/cio.h index bad142c536e1..6d28a62bc67a 100644 --- a/drivers/s390/cio/cio.h +++ b/drivers/s390/cio/cio.h @@ -7,6 +7,7 @@ #include #include #include +#include #include #include #include @@ -49,7 +50,7 @@ struct pmcw { /* Target SCHIB configuration. */ struct schib_config { - u64 mba; + dma64_t mba; u32 intparm; u16 mbi; u32 isc:3; @@ -66,7 +67,7 @@ struct schib_config { struct schib { struct pmcw pmcw; /* path management control word */ union scsw scsw; /* subchannel status word */ - __u64 mba; /* measurement block address */ + dma64_t mba; /* measurement block address */ __u8 mda[4]; /* model dependent area */ } __attribute__ ((packed,aligned(4))); diff --git a/drivers/s390/cio/cmf.c b/drivers/s390/cio/cmf.c index 92ab3d546fe4..66b14fedbd18 100644 --- a/drivers/s390/cio/cmf.c +++ b/drivers/s390/cio/cmf.c @@ -183,7 +183,7 @@ static int set_schib(struct ccw_device *cdev, u32 mme, int mbfc, sch->config.mbfc = mbfc; /* address can be either a block address or a block index */ if (mbfc) - sch->config.mba = address; + sch->config.mba = address ? virt_to_dma64((void *)address) : 0; else sch->config.mbi = address; diff --git a/drivers/s390/cio/device.c b/drivers/s390/cio/device.c index fb591118ecb2..68dd4a62975d 100644 --- a/drivers/s390/cio/device.c +++ b/drivers/s390/cio/device.c @@ -922,7 +922,7 @@ static int ccw_device_move_to_sch(struct ccw_device *cdev, if (!sch_is_pseudo_sch(old_sch)) { spin_lock_irq(&old_sch->lock); - old_enabled = old_sch->schib.pmcw.ena; + old_enabled = old_sch->schib.pmcw.dnv && old_sch->schib.pmcw.ena; rc = 0; if (old_enabled) rc = cio_disable_subchannel(old_sch); @@ -941,7 +941,7 @@ static int ccw_device_move_to_sch(struct ccw_device *cdev, CIO_MSG_EVENT(0, "device_move(0.%x.%04x,0.%x.%04x)=%d\n", cdev->private->dev_id.ssid, cdev->private->dev_id.devno, sch->schid.ssid, - sch->schib.pmcw.dev, rc); + sch->schid.sch_no, rc); if (old_enabled) { /* Try to re-enable the old subchannel. */ spin_lock_irq(&old_sch->lock); @@ -1207,7 +1207,7 @@ static void io_subchannel_quiesce(struct subchannel *sch) cdev = sch_get_cdev(sch); if (cio_is_console(sch->schid)) goto out_unlock; - if (!sch->schib.pmcw.ena) + if (!sch->schib.pmcw.dnv || !sch->schib.pmcw.ena) goto out_unlock; ret = cio_disable_subchannel(sch); if (ret != -EBUSY) @@ -1254,7 +1254,8 @@ static int recovery_check(struct device *dev, void *data) switch (cdev->private->state) { case DEV_STATE_ONLINE: sch = to_subchannel(cdev->dev.parent); - if ((sch->schib.pmcw.pam & sch->opm) == sch->vpm) + if (sch->schib.pmcw.dnv && + (sch->schib.pmcw.pam & sch->opm) == sch->vpm) break; fallthrough; case DEV_STATE_DISCONNECTED: diff --git a/drivers/s390/cio/device_fsm.c b/drivers/s390/cio/device_fsm.c index ab419d40a8a7..b5686c25c83c 100644 --- a/drivers/s390/cio/device_fsm.c +++ b/drivers/s390/cio/device_fsm.c @@ -170,6 +170,9 @@ __recover_lost_chpids(struct subchannel *sch, int old_lpm) int mask, i; struct chp_id chpid; + if (!sch->schib.pmcw.dnv) + return; + chp_id_init(&chpid); for (i = 0; i<8; i++) { mask = 0x80 >> i; diff --git a/drivers/s390/cio/device_ops.c b/drivers/s390/cio/device_ops.c index 61c07b4a0fe8..1f7e83fd5087 100644 --- a/drivers/s390/cio/device_ops.c +++ b/drivers/s390/cio/device_ops.c @@ -142,6 +142,8 @@ int ccw_device_clear(struct ccw_device *cdev, unsigned long intparm) if (!cdev || !cdev->dev.parent) return -ENODEV; sch = to_subchannel(cdev->dev.parent); + if (!sch->schib.pmcw.dnv) + return -ENODEV; if (!sch->schib.pmcw.ena) return -EINVAL; if (cdev->private->state == DEV_STATE_NOT_OPER) @@ -198,6 +200,8 @@ int ccw_device_start_timeout_key(struct ccw_device *cdev, struct ccw1 *cpa, if (!cdev || !cdev->dev.parent) return -ENODEV; sch = to_subchannel(cdev->dev.parent); + if (!sch->schib.pmcw.dnv) + return -ENODEV; if (!sch->schib.pmcw.ena) return -EINVAL; if (cdev->private->state == DEV_STATE_NOT_OPER) @@ -379,6 +383,8 @@ int ccw_device_halt(struct ccw_device *cdev, unsigned long intparm) if (!cdev || !cdev->dev.parent) return -ENODEV; sch = to_subchannel(cdev->dev.parent); + if (!sch->schib.pmcw.dnv) + return -ENODEV; if (!sch->schib.pmcw.ena) return -EINVAL; if (cdev->private->state == DEV_STATE_NOT_OPER) @@ -413,6 +419,8 @@ int ccw_device_resume(struct ccw_device *cdev) if (!cdev || !cdev->dev.parent) return -ENODEV; sch = to_subchannel(cdev->dev.parent); + if (!sch->schib.pmcw.dnv) + return -ENODEV; if (!sch->schib.pmcw.ena) return -EINVAL; if (cdev->private->state == DEV_STATE_NOT_OPER) @@ -482,6 +490,8 @@ struct channel_path_desc_fmt0 *ccw_device_get_chp_desc(struct ccw_device *cdev, struct chp_id chpid; sch = to_subchannel(cdev->dev.parent); + if (!sch->schib.pmcw.dnv) + return NULL; chp_id_init(&chpid); chpid.id = sch->schib.pmcw.chpid[chp_idx]; return chp_get_chp_desc(chpid); @@ -502,9 +512,13 @@ u8 *ccw_device_get_util_str(struct ccw_device *cdev, int chp_idx) struct chp_id chpid; u8 *util_str; + if (!sch->schib.pmcw.dnv) + return NULL; chp_id_init(&chpid); chpid.id = sch->schib.pmcw.chpid[chp_idx]; chp = chpid_to_chp(chpid); + if (!chp) + return NULL; util_str = kmalloc(sizeof(chp->desc_fmt3.util_str), GFP_KERNEL); if (!util_str) @@ -548,6 +562,8 @@ int ccw_device_tm_start_timeout_key(struct ccw_device *cdev, struct tcw *tcw, int rc; sch = to_subchannel(cdev->dev.parent); + if (!sch->schib.pmcw.dnv) + return -ENODEV; if (!sch->schib.pmcw.ena) return -EINVAL; if (cdev->private->state == DEV_STATE_VERIFY) { @@ -652,6 +668,9 @@ int ccw_device_get_mdc(struct ccw_device *cdev, u8 mask) struct chp_id chpid; int mdc = 0, i; + if (!sch->schib.pmcw.dnv) + return 0; + /* Adjust requested path mask to excluded varied off paths. */ if (mask) mask &= sch->lpm; @@ -694,6 +713,8 @@ int ccw_device_tm_intrg(struct ccw_device *cdev) { struct subchannel *sch = to_subchannel(cdev->dev.parent); + if (!sch->schib.pmcw.dnv) + return -ENODEV; if (!sch->schib.pmcw.ena) return -EINVAL; if (cdev->private->state != DEV_STATE_ONLINE) @@ -786,6 +807,8 @@ int ccw_device_get_chpid(struct ccw_device *cdev, int chp_idx, u8 *chpid) if ((chp_idx < 0) || (chp_idx > 7)) return -EINVAL; + if (!sch->schib.pmcw.dnv) + return -ENODEV; mask = 0x80 >> chp_idx; if (!(sch->schib.pmcw.pim & mask)) return -ENODEV; diff --git a/drivers/s390/cio/vfio_ccw_fsm.c b/drivers/s390/cio/vfio_ccw_fsm.c index 5fd94e9d5c61..9a000b0231d6 100644 --- a/drivers/s390/cio/vfio_ccw_fsm.c +++ b/drivers/s390/cio/vfio_ccw_fsm.c @@ -399,7 +399,7 @@ static void fsm_close(struct vfio_ccw_private *private, spin_lock_irq(&sch->lock); - if (!sch->schib.pmcw.ena) + if (!sch->schib.pmcw.dnv || !sch->schib.pmcw.ena) goto err_unlock; ret = cio_disable_subchannel(sch); diff --git a/drivers/s390/crypto/vfio_ap_ops.c b/drivers/s390/crypto/vfio_ap_ops.c index 940c0ff668be..4db878c18f41 100644 --- a/drivers/s390/crypto/vfio_ap_ops.c +++ b/drivers/s390/crypto/vfio_ap_ops.c @@ -277,9 +277,9 @@ static void vfio_ap_free_aqic_resources(struct vfio_ap_queue *q) { if (!q) return; - if (q->saved_isc != VFIO_AP_ISC_INVALID && - !WARN_ON(!(q->matrix_mdev && q->matrix_mdev->kvm))) { - kvm_s390_gisc_unregister(q->matrix_mdev->kvm, q->saved_isc); + if (q->saved_isc != VFIO_AP_ISC_INVALID) { + if (!WARN_ON(!q->matrix_mdev) && q->matrix_mdev->kvm) + kvm_s390_gisc_unregister(q->matrix_mdev->kvm, q->saved_isc); q->saved_isc = VFIO_AP_ISC_INVALID; } if (q->saved_iova && !WARN_ON(!q->matrix_mdev)) { @@ -1935,6 +1935,8 @@ static int apq_status_check(int apqn, struct ap_queue_status *status) * a value indicating a reset needs to be performed again. */ return -EAGAIN; + case AP_RESPONSE_Q_NOT_AVAIL: + return -ENODEV; default: WARN(true, "failed to verify reset of queue %02x.%04x: TAPQ rc=%u\n", @@ -1961,6 +1963,10 @@ static void apq_reset_check(struct work_struct *reset_work) ret = apq_status_check(q->apqn, &status); if (ret == -EIO) return; + if (ret == -ENODEV) { + vfio_ap_free_aqic_resources(q); + return; + } if (ret == -EBUSY) { pr_notice_ratelimited(WAIT_MSG, elapsed, AP_QID_CARD(q->apqn), @@ -2004,6 +2010,7 @@ static void vfio_ap_mdev_reset_queue(struct vfio_ap_queue *q) break; case AP_RESPONSE_DECONFIGURED: case AP_RESPONSE_CHECKSTOPPED: + case AP_RESPONSE_Q_NOT_AVAIL: vfio_ap_free_aqic_resources(q); break; default: @@ -2528,12 +2535,15 @@ void vfio_ap_mdev_remove_queue(struct ap_device *apdev) /* * If the queue is not in the host's AP configuration, then resetting * it will fail with response code 01, (APQN not valid); so, let's make - * sure it is in the host's config. + * sure it is in the host's config. If it is not, free the KVM GISC + * resources. */ if (test_bit_inv(apid, (unsigned long *)matrix_dev->info.apm) && test_bit_inv(apqi, (unsigned long *)matrix_dev->info.aqm)) { vfio_ap_mdev_reset_queue(q); flush_work(&q->reset_work); + } else { + vfio_ap_free_aqic_resources(q); } done: diff --git a/drivers/scsi/libiscsi_tcp.c b/drivers/scsi/libiscsi_tcp.c index 7223bb18b048..d35f93451ee9 100644 --- a/drivers/scsi/libiscsi_tcp.c +++ b/drivers/scsi/libiscsi_tcp.c @@ -480,6 +480,9 @@ static int iscsi_tcp_data_in(struct iscsi_conn *conn, struct iscsi_task *task) int datasn = be32_to_cpu(rhdr->datasn); unsigned total_in_length = task->sc->sdb.length; + if (task->sc->sc_data_direction != DMA_FROM_DEVICE) + return ISCSI_ERR_PROTO; + /* * lib iscsi will update this in the completion handling if there * is status. diff --git a/drivers/scsi/megaraid/megaraid_sas_base.c b/drivers/scsi/megaraid/megaraid_sas_base.c index d83abded2039..f25ed22195fa 100644 --- a/drivers/scsi/megaraid/megaraid_sas_base.c +++ b/drivers/scsi/megaraid/megaraid_sas_base.c @@ -7886,7 +7886,9 @@ megasas_resume(struct device *dev) goto fail_init_mfi; } - if (megasas_get_ctrl_info(instance) != DCMD_SUCCESS) + scoped_guard(mutex, &instance->reset_mutex) + rval = megasas_get_ctrl_info(instance); + if (rval != DCMD_SUCCESS) goto fail_init_mfi; tasklet_init(&instance->isr_tasklet, instance->instancet->tasklet, diff --git a/drivers/scsi/sd_zbc.c b/drivers/scsi/sd_zbc.c index 56e455fb5add..456beaf2e769 100644 --- a/drivers/scsi/sd_zbc.c +++ b/drivers/scsi/sd_zbc.c @@ -589,7 +589,7 @@ int sd_zbc_revalidate_zones(struct scsi_disk *sdkp) int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim, u8 buf[SD_BUF_SIZE]) { - unsigned int nr_zones; + u64 nr_zones; u32 zone_blocks = 0; int ret; @@ -621,6 +621,12 @@ int sd_zbc_read_zones(struct scsi_disk *sdkp, struct queue_limits *lim, goto err; nr_zones = round_up(sdkp->capacity, zone_blocks) >> ilog2(zone_blocks); + if (nr_zones > INT_MAX) { + sd_printk(KERN_ERR, sdkp, "Too many zones (%llu)\n", + nr_zones); + ret = -EINVAL; + goto err; + } sdkp->early_zone_info.nr_zones = nr_zones; sdkp->early_zone_info.zone_blocks = zone_blocks; diff --git a/drivers/thermal/gov_step_wise.c b/drivers/thermal/gov_step_wise.c index ea277c466d8d..4fa4377f0d42 100644 --- a/drivers/thermal/gov_step_wise.c +++ b/drivers/thermal/gov_step_wise.c @@ -65,14 +65,12 @@ static unsigned long get_target_state(struct thermal_instance *instance, min(instance->lower + 1, instance->upper), instance->upper); } else if (trend == THERMAL_TREND_DROPPING) { - if (cur_state <= instance->lower) - return THERMAL_NO_TARGET; - /* - * If 'throttle' is false, no mitigation is necessary, so - * request the lower state for this instance. + * If 'throttle' is false, no mitigation is necessary and + * passive polling is already deactivated, so clear this + * instance state by returning THERMAL_NO_TARGET. */ - return instance->lower; + return THERMAL_NO_TARGET; } return instance->target; diff --git a/drivers/ufs/core/ufshcd.c b/drivers/ufs/core/ufshcd.c index 86f49c0522ec..bc33da89af26 100644 --- a/drivers/ufs/core/ufshcd.c +++ b/drivers/ufs/core/ufshcd.c @@ -6748,11 +6748,17 @@ static void ufshcd_err_handling_prepare(struct ufs_hba *hba) } /* Wait for ongoing ufshcd_queuecommand() calls to finish. */ blk_mq_quiesce_tagset(&hba->host->tag_set); + /* + * Internal commands are submitted on the pseudo SCSI device. Let them + * through so that the error handler can recover the link. + */ + blk_mq_unquiesce_queue(hba->host->pseudo_sdev->request_queue); cancel_work_sync(&hba->eeh_work); } static void ufshcd_err_handling_unprepare(struct ufs_hba *hba) { + blk_mq_quiesce_queue_nowait(hba->host->pseudo_sdev->request_queue); blk_mq_unquiesce_tagset(&hba->host->tag_set); ufshcd_release(hba); if (ufshcd_is_clkscaling_supported(hba)) diff --git a/drivers/ufs/host/ufshcd-pltfrm.c b/drivers/ufs/host/ufshcd-pltfrm.c index c2dafb583cf5..ac264786d2a6 100644 --- a/drivers/ufs/host/ufshcd-pltfrm.c +++ b/drivers/ufs/host/ufshcd-pltfrm.c @@ -206,7 +206,11 @@ static void ufshcd_init_lanes_per_dir(struct ufs_hba *hba) dev_dbg(hba->dev, "%s: failed to read lanes-per-direction, ret=%d\n", __func__, ret); - hba->lanes_per_direction = UFSHCD_DEFAULT_LANES_PER_DIRECTION; + /* Old R-Car S4 DTBs lack "lanes-per-direction = <1>" */ + if (of_device_is_compatible(dev->of_node, "renesas,r8a779f0-ufs")) + hba->lanes_per_direction = 1; + else + hba->lanes_per_direction = UFSHCD_DEFAULT_LANES_PER_DIRECTION; } } diff --git a/fs/autofs/inode.c b/fs/autofs/inode.c index 6b15a3717ba7..066c16f2ea56 100644 --- a/fs/autofs/inode.c +++ b/fs/autofs/inode.c @@ -51,6 +51,10 @@ void autofs_kill_sb(struct super_block *sb) if (sbi) { /* Free wait queues, close pipe */ autofs_catatonic_mode(sbi); + if (sbi->pipe) { + fput(sbi->pipe); + sbi->pipe = NULL; + } put_pid(sbi->oz_pgrp); } diff --git a/fs/bpf_fs_kfuncs.c b/fs/bpf_fs_kfuncs.c index f1863a891db6..d77604a8996b 100644 --- a/fs/bpf_fs_kfuncs.c +++ b/fs/bpf_fs_kfuncs.c @@ -419,10 +419,6 @@ BTF_ID(func, bpf_lsm_inode_rmdir) BTF_ID(func, bpf_lsm_inode_setattr) BTF_ID(func, bpf_lsm_inode_setxattr) BTF_ID(func, bpf_lsm_inode_unlink) -#ifdef CONFIG_SECURITY_PATH -BTF_ID(func, bpf_lsm_path_unlink) -BTF_ID(func, bpf_lsm_path_rmdir) -#endif /* CONFIG_SECURITY_PATH */ BTF_SET_END(d_inode_locked_hooks) bool bpf_lsm_has_d_inode_locked(const struct bpf_prog *prog) diff --git a/fs/fs-writeback.c b/fs/fs-writeback.c index fdb8766d275a..c0363c3bdb0e 100644 --- a/fs/fs-writeback.c +++ b/fs/fs-writeback.c @@ -726,19 +726,34 @@ static bool isw_prepare_wbs_switch(struct bdi_writeback *new_wb, struct inode_switch_wbs_context *isw, struct list_head *list, int *nr) { - struct inode *inode; + struct inode *inode, *tmp; + LIST_HEAD(scanned); + bool full = false; + + /* + * Walk from the oldest end and move scanned inodes to the newest + * end, so the next scan resumes at unscanned inodes instead of + * re-walking an ever-growing run of prepared and skipped ones. + * For b_dirty_time this keeps the oldest unscanned inode at the + * end move_expired_inodes() picks from; b_attached is unordered. + */ + list_for_each_entry_safe_reverse(inode, tmp, list, i_io_list) { + list_move(&inode->i_io_list, &scanned); - list_for_each_entry(inode, list, i_io_list) { if (!inode_prepare_wbs_switch(inode, new_wb)) continue; isw->inodes[*nr] = inode; (*nr)++; - if (*nr >= WB_MAX_INODES_PER_ISW - 1) - return true; + if (*nr >= WB_MAX_INODES_PER_ISW - 1) { + full = true; + break; + } } - return false; + list_splice(&scanned, list); + + return full; } /** diff --git a/fs/inode.c b/fs/inode.c index 95e981b4e19c..c6b08e8c5f62 100644 --- a/fs/inode.c +++ b/fs/inode.c @@ -883,7 +883,6 @@ void evict_inodes(struct super_block *sb) struct inode *inode; LIST_HEAD(dispose); -again: spin_lock(&sb->s_inode_list_lock); list_for_each_entry(inode, &sb->s_inodes, i_sb_list) { if (icount_read_once(inode)) @@ -902,19 +901,19 @@ void evict_inodes(struct super_block *sb) inode_state_set(inode, I_FREEING); inode_lru_list_del(inode); spin_unlock(&inode->i_lock); - list_add(&inode->i_lru, &dispose); /* - * We can have a ton of inodes to evict at unmount time given - * enough memory, check to see if we need to go to sleep for a - * bit so we don't livelock. + * Keep this inode out of dispose so it stays on s_inodes while + * the list lock is dropped. I_FREEING prevents new references + * and leaves eviction to us, so we can resume the walk from it. */ if (need_resched()) { spin_unlock(&sb->s_inode_list_lock); cond_resched(); dispose_list(&dispose); - goto again; + spin_lock(&sb->s_inode_list_lock); } + list_add(&inode->i_lru, &dispose); } spin_unlock(&sb->s_inode_list_lock); diff --git a/fs/kernfs/mount.c b/fs/kernfs/mount.c index f183a96778b9..a57399021c8b 100644 --- a/fs/kernfs/mount.c +++ b/fs/kernfs/mount.c @@ -434,8 +434,8 @@ void kernfs_kill_sb(struct super_block *sb) up_write(&root->kernfs_supers_rwsem); /* - * Remove the superblock from fs_supers/s_instances - * so we can't find it, before freeing kernfs_super_info. + * Mark the superblock dead so sget_fc() can't find it, + * before freeing kernfs_super_info. */ kill_anon_super(sb); kfree(info); diff --git a/fs/netfs/buffered_read.c b/fs/netfs/buffered_read.c index 424df70a5c30..105194de6e13 100644 --- a/fs/netfs/buffered_read.c +++ b/fs/netfs/buffered_read.c @@ -482,15 +482,14 @@ static int netfs_read_gaps(struct file *file, struct folio *folio) struct netfs_group *group = netfs_folio_group(folio); struct netfs_folio *finfo = netfs_folio_info(folio); struct netfs_inode *ctx = netfs_inode(mapping->host); - struct folio *sink = NULL; - struct bio_vec *bvec; + struct bio_vec *bvec = NULL; unsigned int from = finfo->dirty_offset; unsigned int to = from + finfo->dirty_len; - unsigned int off = 0, i = 0; + unsigned int off = 0; size_t flen = folio_size(folio); size_t nr_bvec = flen / PAGE_SIZE + 2; size_t part; - int ret; + int ret, i = 0, sink_from = -1, sink_to = -1; _enter("%lx", folio->index); @@ -515,24 +514,23 @@ static int netfs_read_gaps(struct file *file, struct folio *folio) if (!bvec) goto discard; - sink = folio_alloc(GFP_KERNEL, 0); - if (!sink) { - kfree(bvec); - goto discard; - } - trace_netfs_folio(folio, netfs_folio_trace_read_gaps); - rreq->direct_bv = bvec; - rreq->direct_bv_count = nr_bvec; if (from > 0) { bvec_set_folio(&bvec[i++], folio, from, 0); off = from; } + sink_from = i; while (off < to) { + struct folio *sink = folio_alloc(GFP_KERNEL, 0); + + if (!sink) + goto discard; part = min_t(size_t, to - off, PAGE_SIZE); - bvec_set_folio(&bvec[i++], sink, part, 0); + bvec_set_folio(&bvec[i], sink, part, 0); off += part; + sink_to = i; + i++; } if (to < flen) bvec_set_folio(&bvec[i++], folio, flen - to, to); @@ -553,8 +551,10 @@ static int netfs_read_gaps(struct file *file, struct folio *folio) folio_mark_uptodate(folio); } - if (sink) - folio_put(sink); + if (sink_to >= 0) + for (; sink_from <= sink_to; sink_from++) + folio_put(bvec_folio(&bvec[sink_from])); + kfree(bvec); folio_unlock(folio); netfs_put_request(rreq, netfs_rreq_trace_put_return); return ret < 0 ? ret : 0; @@ -563,6 +563,10 @@ static int netfs_read_gaps(struct file *file, struct folio *folio) netfs_put_failed_request(rreq); alloc_error: folio_unlock(folio); + if (sink_to >= 0) + for (; sink_from <= sink_to; sink_from++) + folio_put(bvec_folio(&bvec[sink_from])); + kfree(bvec); return ret; } diff --git a/fs/netfs/objects.c b/fs/netfs/objects.c index 7f6a3e912602..ad549daa9c79 100644 --- a/fs/netfs/objects.c +++ b/fs/netfs/objects.c @@ -34,7 +34,7 @@ struct netfs_io_request *netfs_alloc_request(struct address_space *mapping, rreq = mempool_alloc(mempool, gfp); } else { - rreq = mempool->alloc(gfp, mempool->pool_data); + rreq = mempool_alloc_noreserve(mempool, gfp); if (!rreq) return ERR_PTR(-ENOMEM); } @@ -214,7 +214,7 @@ struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq struct kmem_cache *cache = mempool->pool_data; if (rreq->gfp == GFP_KERNEL) - subreq = mempool->alloc(rreq->gfp, mempool->pool_data); + subreq = mempool_alloc_noreserve(mempool, rreq->gfp); else subreq = mempool_alloc(mempool, rreq->gfp); if (!subreq) diff --git a/fs/netfs/read_collect.c b/fs/netfs/read_collect.c index 5cf22087d243..a94197ef0181 100644 --- a/fs/netfs/read_collect.c +++ b/fs/netfs/read_collect.c @@ -435,6 +435,11 @@ static void netfs_rreq_assess_single(struct netfs_io_request *rreq) netfs_single_mark_inode_dirty(rreq->inode); } + /* To do DIO, the cache has to round the size up, so we need to undo + * the rounding. + */ + rreq->transferred = min(rreq->transferred, rreq->i_size); + if (rreq->iocb) { rreq->iocb->ki_pos += rreq->transferred; if (rreq->iocb->ki_complete) { diff --git a/fs/netfs/rolling_buffer.c b/fs/netfs/rolling_buffer.c index 424e77a9a109..d30d5ef6d86e 100644 --- a/fs/netfs/rolling_buffer.c +++ b/fs/netfs/rolling_buffer.c @@ -29,7 +29,7 @@ struct folio_queue *netfs_folioq_alloc(unsigned int rreq_id, gfp_t gfp, struct folio_queue *fq; if (gfp == GFP_KERNEL) - fq = netfs_folioq_pool.alloc(gfp, netfs_folioq_pool.pool_data); + fq = mempool_alloc_noreserve(&netfs_folioq_pool, gfp); else fq = mempool_alloc(&netfs_folioq_pool, gfp); if (fq) { diff --git a/fs/ntfs3/inode.c b/fs/ntfs3/inode.c index 0c9bd669117d..ea4b9193b07d 100644 --- a/fs/ntfs3/inode.c +++ b/fs/ntfs3/inode.c @@ -1648,10 +1648,10 @@ int ntfs_create_inode(struct mnt_idmap *idmap, struct inode *dir, goto out6; /* - * Call 'd_instantiate' after inode->i_op is set + * Call 'd_instantiate_new' after inode->i_op is set * but before finish_open. */ - d_instantiate(dentry, inode); + d_instantiate_new(dentry, inode); /* Set original time. inode times (i_ctime) may be changed in ntfs_init_acl. */ inode_set_atime_to_ts(inode, ni->i_crtime); @@ -1699,9 +1699,6 @@ int ntfs_create_inode(struct mnt_idmap *idmap, struct inode *dir, if (!fnd) ni_unlock(dir_ni); - if (!err) - unlock_new_inode(inode); - return err; } diff --git a/fs/ocfs2/namei.c b/fs/ocfs2/namei.c index ea37a5058089..4e9f8dd63dcb 100644 --- a/fs/ocfs2/namei.c +++ b/fs/ocfs2/namei.c @@ -336,13 +336,8 @@ static int ocfs2_mknod(struct mnt_idmap *idmap, goto leave; /* calculate meta data/clusters for setting security and acl xattr */ - status = ocfs2_calc_xattr_init(dir, mode, &si, &want_clusters, - &xattr_credits, &want_meta, - &acl_state); - if (status < 0) { - mlog_errno(status); - goto leave; - } + ocfs2_calc_xattr_init(dir, mode, &si, &want_clusters, &xattr_credits, + &want_meta, &acl_state); /* Reserve a cluster if creating an extent based directory. */ if (S_ISDIR(mode) && !ocfs2_supports_inline_data(osb)) { diff --git a/fs/ocfs2/xattr.c b/fs/ocfs2/xattr.c index 35bcbb0ff607..bfafe059bedf 100644 --- a/fs/ocfs2/xattr.c +++ b/fs/ocfs2/xattr.c @@ -635,12 +635,11 @@ int ocfs2_calc_security_init(struct inode *dir, return ret; } -int ocfs2_calc_xattr_init(struct inode *dir, umode_t mode, - struct ocfs2_security_xattr_info *si, - int *want_clusters, int *xattr_credits, - int *want_meta, struct ocfs2_acl_state *acl_state) +void ocfs2_calc_xattr_init(struct inode *dir, umode_t mode, + struct ocfs2_security_xattr_info *si, + int *want_clusters, int *xattr_credits, + int *want_meta, struct ocfs2_acl_state *acl_state) { - int ret = 0; struct ocfs2_super *osb = OCFS2_SB(dir->i_sb); int s_size = 0, a_size = 0, acl_len = 0, new_clusters; @@ -662,7 +661,7 @@ int ocfs2_calc_xattr_init(struct inode *dir, umode_t mode, } if (!(s_size + a_size)) - return ret; + return; /* * The max space of security xattr taken inline is @@ -728,8 +727,6 @@ int ocfs2_calc_xattr_init(struct inode *dir, umode_t mode, } } } - - return ret; } static int ocfs2_xattr_extend_allocation(struct inode *inode, diff --git a/fs/ocfs2/xattr.h b/fs/ocfs2/xattr.h index 5e18513277f1..887cc1a18b1a 100644 --- a/fs/ocfs2/xattr.h +++ b/fs/ocfs2/xattr.h @@ -59,10 +59,10 @@ int ocfs2_calc_security_init(struct inode *, int *, int *, struct ocfs2_alloc_context **); struct ocfs2_acl_state; -int ocfs2_calc_xattr_init(struct inode *dir, umode_t mode, - struct ocfs2_security_xattr_info *si, - int *want_clusters, int *xattr_credits, - int *want_meta, struct ocfs2_acl_state *acl_state); +void ocfs2_calc_xattr_init(struct inode *dir, umode_t mode, + struct ocfs2_security_xattr_info *si, + int *want_clusters, int *xattr_credits, + int *want_meta, struct ocfs2_acl_state *acl_state); /* * xattrs can live inside an inode, as part of an external xattr block, diff --git a/fs/overlayfs/overlayfs.h b/fs/overlayfs/overlayfs.h index b75df37f70ac..3fea90fe37c7 100644 --- a/fs/overlayfs/overlayfs.h +++ b/fs/overlayfs/overlayfs.h @@ -254,8 +254,10 @@ static inline struct dentry *ovl_do_mkdir(struct ovl_fs *ofs, { struct dentry *ret; + /* vfs_mkdir() drops @dentry on failure and may replace it on success */ + pr_debug("mkdir(%pd2, 0%o)\n", dentry, mode); ret = vfs_mkdir(ovl_upper_mnt_idmap(ofs), dir, dentry, mode, NULL); - pr_debug("mkdir(%pd2, 0%o) = %i\n", dentry, mode, PTR_ERR_OR_ZERO(ret)); + pr_debug("...mkdir = %i\n", PTR_ERR_OR_ZERO(ret)); return ret; } diff --git a/fs/smb/client/cached_dir.c b/fs/smb/client/cached_dir.c index 88d5e9a32f28..647fa26da4d2 100644 --- a/fs/smb/client/cached_dir.c +++ b/fs/smb/client/cached_dir.c @@ -8,6 +8,7 @@ #include #include "cifsglob.h" #include "cifsproto.h" +#include "../common/smb2status.h" #include "cifs_debug.h" #include "smb2proto.h" #include "cached_dir.h" @@ -323,25 +324,37 @@ int open_cached_dir(unsigned int xid, struct cifs_tcon *tcon, rc = compound_send_recv(xid, ses, server, flags, 2, rqst, resp_buftype, rsp_iov); - if (rc) { - if (rc == -EREMCHG) { - tcon->need_reconnect = true; - pr_warn_once("server share %s deleted\n", - tcon->tree_name); - } - goto oshr_free; + if (rc == -EREMCHG) { + tcon->need_reconnect = true; + pr_warn_once("server share %s deleted\n", + tcon->tree_name); } - cfid->is_open = true; - spin_lock(&cfids->cfid_list_lock); + if (!rsp_iov[0].iov_base || rsp_iov[0].iov_len < sizeof(*o_rsp)) { + if (!rc) + rc = -EIO; + goto oshr_free; + } o_rsp = (struct smb2_create_rsp *)rsp_iov[0].iov_base; + if (o_rsp->hdr.Status != STATUS_SUCCESS) { + if (!rc) + rc = -EIO; + goto oshr_free; + } + oparms.fid->persistent_fid = o_rsp->PersistentFileId; oparms.fid->volatile_fid = o_rsp->VolatileFileId; #ifdef CONFIG_CIFS_DEBUG2 oparms.fid->mid = le64_to_cpu(o_rsp->hdr.MessageId); #endif /* CIFS_DEBUG2 */ + cfid->is_open = true; + atomic_inc(&tcon->num_remote_opens); + if (rc) + goto oshr_free; + + spin_lock(&cfids->cfid_list_lock); if (o_rsp->OplockLevel != SMB2_OPLOCK_LEVEL_LEASE) { spin_unlock(&cfids->cfid_list_lock); @@ -408,7 +421,6 @@ int open_cached_dir(unsigned int xid, struct cifs_tcon *tcon, close_cached_dir(cfid); } else { *ret_cfid = cfid; - atomic_inc(&tcon->num_remote_opens); } kfree(utf16_path); diff --git a/fs/smb/client/dir.c b/fs/smb/client/dir.c index b0ddcaa2d815..19e6b1a6b9b0 100644 --- a/fs/smb/client/dir.c +++ b/fs/smb/client/dir.c @@ -203,7 +203,7 @@ static int __cifs_do_create(struct inode *dir, struct dentry *direntry, struct tcon_link *tlink, unsigned int oflags, umode_t mode, __u32 *oplock, struct cifs_fid *fid, struct cifs_open_info_data *buf, - struct inode **inode) + struct inode **inode, bool *opened) { int rc = -ENOENT; int create_options = CREATE_NOT_DIR; @@ -220,6 +220,7 @@ static int __cifs_do_create(struct inode *dir, struct dentry *direntry, __le32 lease_flags = 0; *inode = NULL; + *opened = false; *oplock = 0; if (tcon->ses->server->oplocks) *oplock = REQ_OPLOCK; @@ -236,6 +237,7 @@ static int __cifs_do_create(struct inode *dir, struct dentry *direntry, oflags, oplock, &fid->netfid, xid); switch (rc) { case 0: + *opened = true; if (newinode == NULL) { /* query inode info */ goto cifs_create_get_file_info; @@ -257,11 +259,9 @@ static int __cifs_do_create(struct inode *dir, struct dentry *direntry, /* * The server may allow us to open things like * FIFOs, but the client isn't set up to deal - * with that. If it's not a regular file, just - * close it and proceed as if it were a normal - * lookup. + * with that. Keep the handle until the caller + * can finish the lookup. */ - CIFSSMBClose(xid, tcon, fid->netfid); goto cifs_create_get_file_info; } /* success, no need to query */ @@ -388,6 +388,7 @@ static int __cifs_do_create(struct inode *dir, struct dentry *direntry, } return rc; } + *opened = true; if (rdwr_for_fscache == 2) cifs_invalidate_cache(dir, FSCACHE_INVAL_DIO_WRITE); @@ -479,7 +480,7 @@ static int __cifs_do_create(struct inode *dir, struct dentry *direntry, return rc; out_err: - if (server->ops->close) + if (*opened && server->ops->close) server->ops->close(xid, tcon, fid); if (newinode) iput(newinode); @@ -491,7 +492,7 @@ static int cifs_do_create(struct inode *dir, struct dentry *direntry, unsigned int oflags, umode_t mode, __u32 *oplock, struct cifs_fid *fid, struct cifs_open_info_data *buf, - struct inode **inode) + struct inode **inode, bool *opened) { void *page = alloc_dentry_path(); const char *full_path; @@ -500,10 +501,11 @@ static int cifs_do_create(struct inode *dir, struct dentry *direntry, full_path = build_path_from_dentry(direntry, page); if (IS_ERR(full_path)) { rc = PTR_ERR(full_path); + *opened = false; } else { rc = __cifs_do_create(dir, direntry, full_path, xid, tlink, oflags, mode, oplock, - fid, buf, inode); + fid, buf, inode, opened); } free_dentry_path(page); return rc; @@ -533,6 +535,8 @@ int cifs_atomic_open(struct inode *dir, struct dentry *direntry, struct inode *inode; unsigned int xid; __u32 oplock; + bool is_regular; + bool opened; int rc; if (unlikely(cifs_forced_shutdown(cifs_sb))) @@ -585,12 +589,26 @@ int cifs_atomic_open(struct inode *dir, struct dentry *direntry, cifs_add_pending_open(&fid, tlink, &open); rc = cifs_do_create(dir, direntry, xid, tlink, oflags, mode, - &oplock, &fid, &buf, &inode); + &oplock, &fid, &buf, &inode, &opened); if (rc) { cifs_del_pending_open(&open); goto out; } + is_regular = S_ISREG(inode->i_mode); + if (!is_regular || !opened) { + if (opened && server->ops->close) + server->ops->close(xid, tcon, &fid); + cifs_del_pending_open(&open); + if (S_ISLNK(inode->i_mode) && + (oflags & (O_NOFOLLOW | __O_REGULAR)) == + (O_NOFOLLOW | __O_REGULAR) && !(oflags & O_EXCL)) { + iput(inode); + rc = -ELOOP; + goto out; + } + } + if (d_in_lookup(direntry)) { alias = d_splice_alias(inode, direntry); if (!IS_ERR_OR_NULL(alias)) @@ -599,9 +617,15 @@ int cifs_atomic_open(struct inode *dir, struct dentry *direntry, d_instantiate(direntry, inode); } - if ((oflags & (O_CREAT | O_EXCL)) == (O_CREAT | O_EXCL)) + if (is_regular && opened && + (oflags & (O_CREAT | O_EXCL)) == (O_CREAT | O_EXCL)) file->f_mode |= FMODE_CREATED; + if (!is_regular || !opened) { + rc = finish_no_open(file, NULL); + goto out; + } + rc = finish_open(file, direntry, generic_file_open); if (rc) { if (server->ops->close) @@ -664,6 +688,7 @@ int cifs_create(struct mnt_idmap *idmap, struct inode *dir, struct inode *inode; struct cifs_fid fid; __u32 oplock; + bool opened; struct cifs_open_info_data buf = {}; cifs_dbg(FYI, "cifs_create parent inode = 0x%p name is: %pd and dentry = 0x%p\n", @@ -686,10 +711,10 @@ int cifs_create(struct mnt_idmap *idmap, struct inode *dir, server->ops->new_lease_key(&fid); rc = cifs_do_create(dir, direntry, xid, tlink, oflags, - mode, &oplock, &fid, &buf, &inode); + mode, &oplock, &fid, &buf, &inode, &opened); if (!rc) { d_instantiate(direntry, inode); - if (server->ops->close) + if (opened && server->ops->close) server->ops->close(xid, tcon, &fid); } @@ -1082,6 +1107,7 @@ int cifs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, struct inode *inode; unsigned int xid; __u32 oplock; + bool opened; int namelen; int rc; @@ -1120,7 +1146,7 @@ int cifs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, namelen = scnprintf(name, namesize, CIFS_TMPNAME_PREFIX "%x", atomic_inc_return(&cifs_tmpcounter)); rc = __cifs_do_create(dir, dentry, path, xid, tlink, oflags, - mode, &oplock, &fid, NULL, &inode); + mode, &oplock, &fid, NULL, &inode, &opened); if (!rc) { rc = d_mark_tmpfile_name(file, &QSTR_LEN(name, namelen)); if (rc) { diff --git a/fs/smb/client/smb2inode.c b/fs/smb/client/smb2inode.c index 441f107ed202..970503a7674e 100644 --- a/fs/smb/client/smb2inode.c +++ b/fs/smb/client/smb2inode.c @@ -596,8 +596,10 @@ static int smb2_compound_op(const unsigned int xid, struct cifs_tcon *tcon, /* smb2_parse_contexts() fills idata->fi.IndexNumber */ rc = smb2_parse_contexts(server, &rsp_iov[0], &oparms->fid->epoch, oparms->fid->lease_key, &oplock, &idata->fi, NULL); - if (rc) + if (rc) { cifs_dbg(VFS, "rc: %d parsing context of compound op\n", rc); + tmp_rc = rc; + } } for (i = 0; i < num_cmds; i++) { @@ -1118,7 +1120,7 @@ smb2_unlink(const unsigned int xid, struct cifs_tcon *tcon, const char *name, struct kvec close_iov; int resp_buftype[2]; struct cifs_fid fid; - int flags = 0; + int flags = CIFS_CP_CREATE_CLOSE_OP; __u8 oplock; int rc; diff --git a/fs/smb/client/smb2misc.c b/fs/smb/client/smb2misc.c index 9068175e57cd..596388acb31e 100644 --- a/fs/smb/client/smb2misc.c +++ b/fs/smb/client/smb2misc.c @@ -821,7 +821,8 @@ smb2_cancelled_close_fid(struct work_struct *work) */ static int __smb2_handle_cancelled_cmd(struct cifs_tcon *tcon, __u16 cmd, __u64 mid, - __u64 persistent_fid, __u64 volatile_fid) + __u64 persistent_fid, __u64 volatile_fid, + bool account_remote_open) { struct close_cancelled_open *cancelled; @@ -835,6 +836,8 @@ __smb2_handle_cancelled_cmd(struct cifs_tcon *tcon, __u16 cmd, __u64 mid, cancelled->cmd = cmd; cancelled->mid = mid; INIT_WORK(&cancelled->work, smb2_cancelled_close_fid); + if (account_remote_open) + atomic_inc(&tcon->num_remote_opens); WARN_ON(queue_work(cifsiod_wq, &cancelled->work) == false); return 0; @@ -871,7 +874,7 @@ smb2_handle_cancelled_close(struct cifs_tcon *tcon, __u64 persistent_fid, spin_unlock(&tcon->tc_lock); rc = __smb2_handle_cancelled_cmd(tcon, SMB2_CLOSE_HE, 0, - persistent_fid, volatile_fid); + persistent_fid, volatile_fid, false); if (rc) cifs_put_tcon(tcon, netfs_trace_tcon_ref_put_cancelled_close); @@ -899,7 +902,7 @@ smb2_handle_cancelled_mid(struct mid_q_entry *mid, struct TCP_Server_Info *serve le16_to_cpu(hdr->Command), le64_to_cpu(hdr->MessageId), rsp->PersistentFileId, - rsp->VolatileFileId); + rsp->VolatileFileId, true); if (rc) cifs_put_tcon(tcon, netfs_trace_tcon_ref_put_cancelled_mid); diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c index 0c8478d1d698..46507f8bc68b 100644 --- a/fs/smb/client/smb2ops.c +++ b/fs/smb/client/smb2ops.c @@ -4569,25 +4569,37 @@ smb3_create_lease_buf(u8 *lease_key, u8 oplock, u8 *parent_lease_key, __le32 fla static __u8 smb2_parse_lease_buf(void *buf, __u16 *epoch, char *lease_key) { - struct create_lease *lc = (struct create_lease *)buf; + struct create_context *cc = buf; + struct lease_context lc; *epoch = 0; /* not used */ - if (lc->lcontext.LeaseFlags & SMB2_LEASE_FLAG_BREAK_IN_PROGRESS_LE) + if (le32_to_cpu(cc->DataLength) != sizeof(lc)) + return 0; + + memcpy(&lc, (u8 *)cc + le16_to_cpu(cc->DataOffset), sizeof(lc)); + if (lc.LeaseFlags & SMB2_LEASE_FLAG_BREAK_IN_PROGRESS_LE) return SMB2_OPLOCK_LEVEL_NOCHANGE; - return le32_to_cpu(lc->lcontext.LeaseState); + return le32_to_cpu(lc.LeaseState); } static __u8 smb3_parse_lease_buf(void *buf, __u16 *epoch, char *lease_key) { - struct create_lease_v2 *lc = (struct create_lease_v2 *)buf; + struct create_context *cc = buf; + struct lease_context_v2 lc; + + if (le32_to_cpu(cc->DataLength) != sizeof(lc)) { + *epoch = 0; + return 0; + } - *epoch = le16_to_cpu(lc->lcontext.Epoch); - if (lc->lcontext.LeaseFlags & SMB2_LEASE_FLAG_BREAK_IN_PROGRESS_LE) + memcpy(&lc, (u8 *)cc + le16_to_cpu(cc->DataOffset), sizeof(lc)); + *epoch = le16_to_cpu(lc.Epoch); + if (lc.LeaseFlags & SMB2_LEASE_FLAG_BREAK_IN_PROGRESS_LE) return SMB2_OPLOCK_LEVEL_NOCHANGE; if (lease_key) - memcpy(lease_key, &lc->lcontext.LeaseKey, SMB2_LEASE_KEY_SIZE); - return le32_to_cpu(lc->lcontext.LeaseState); + memcpy(lease_key, lc.LeaseKey, SMB2_LEASE_KEY_SIZE); + return le32_to_cpu(lc.LeaseState); } static unsigned int diff --git a/fs/smb/client/smb2pdu.c b/fs/smb/client/smb2pdu.c index 880ce12f50c4..3d7ead36d1a0 100644 --- a/fs/smb/client/smb2pdu.c +++ b/fs/smb/client/smb2pdu.c @@ -2379,23 +2379,32 @@ create_reconnect_durable_buf(struct cifs_fid *fid) static void parse_query_id_ctxt(struct create_context *cc, struct smb2_file_all_info *buf) { - struct create_disk_id_rsp *pdisk_id = (struct create_disk_id_rsp *)cc; + u16 doff = le16_to_cpu(cc->DataOffset); + u32 dlen = le32_to_cpu(cc->DataLength); + u8 *beg; - cifs_dbg(FYI, "parse query id context 0x%llx 0x%llx\n", - pdisk_id->DiskFileId, pdisk_id->VolumeId); - buf->IndexNumber = pdisk_id->DiskFileId; + if (dlen < sizeof(__le64)) + return; + + beg = (u8 *)cc + doff; + memcpy(&buf->IndexNumber, beg, sizeof(__le64)); + cifs_dbg(FYI, "parse query id context 0x%llx\n", + le64_to_cpu(buf->IndexNumber)); } static void parse_posix_ctxt(struct create_context *cc, struct smb2_file_all_info *info, struct create_posix_rsp *posix) { - int sid_len; u8 *beg = (u8 *)cc + le16_to_cpu(cc->DataOffset); - u8 *end = beg + le32_to_cpu(cc->DataLength); + u32 dlen = le32_to_cpu(cc->DataLength); + u8 *end = beg + dlen; + int sid_len; u8 *sid; memset(posix, 0, sizeof(*posix)); + if (dlen < 3 * sizeof(__le32)) + return; posix->nlink = get_unaligned_le32(beg); posix->reparse_tag = get_unaligned_le32(beg + 4); @@ -2431,6 +2440,7 @@ int smb2_parse_contexts(struct TCP_Server_Info *server, struct smb2_create_rsp *rsp = rsp_iov->iov_base; struct create_context *cc; size_t rem, off, len; + size_t cc_len; size_t doff, dlen; size_t noff, nlen; char *name; @@ -2453,29 +2463,41 @@ int smb2_parse_contexts(struct TCP_Server_Info *server, buf->IndexNumber = 0; while (rem >= sizeof(*cc)) { + off = le32_to_cpu(cc->Next); + if (off) { + if ((off & 0x7) || off >= rem || off < sizeof(*cc)) + return -EINVAL; + cc_len = off; + } else { + cc_len = rem; + } + doff = le16_to_cpu(cc->DataOffset); dlen = le32_to_cpu(cc->DataLength); - if (check_add_overflow(doff, dlen, &len) || len > rem) + if (doff < sizeof(*cc) || + check_add_overflow(doff, dlen, &len) || len > cc_len) return -EINVAL; noff = le16_to_cpu(cc->NameOffset); nlen = le16_to_cpu(cc->NameLength); - if (noff + nlen > doff) + if (noff < sizeof(*cc) || + check_add_overflow(noff, nlen, &len) || len > cc_len || + (dlen && len > doff)) return -EINVAL; name = (char *)cc + noff; switch (nlen) { case 4: - if (!strncmp(name, SMB2_CREATE_REQUEST_LEASE, 4)) { + if (dlen && !strncmp(name, SMB2_CREATE_REQUEST_LEASE, 4)) { *oplock = server->ops->parse_lease_buf(cc, epoch, lease_key); - } else if (buf && + } else if (dlen && buf && !strncmp(name, SMB2_CREATE_QUERY_ON_DISK_ID, 4)) { parse_query_id_ctxt(cc, buf); } break; case 16: - if (posix && !memcmp(name, smb3_create_tag_posix, 16)) + if (dlen && posix && !memcmp(name, smb3_create_tag_posix, 16)) parse_posix_ctxt(cc, buf, posix); break; default: @@ -2487,13 +2509,18 @@ int smb2_parse_contexts(struct TCP_Server_Info *server, } off = le32_to_cpu(cc->Next); - if (!off) + if (!off) { + rem = 0; break; + } if (check_sub_overflow(rem, off, &rem)) return -EINVAL; cc = (struct create_context *)((u8 *)cc + off); } + if (rem) + return -EINVAL; + if (rsp->OplockLevel != SMB2_OPLOCK_LEVEL_LEASE) *oplock = rsp->OplockLevel; @@ -3389,6 +3416,9 @@ SMB2_open(const unsigned int xid, struct cifs_open_parms *oparms, __le16 *path, rc = smb2_parse_contexts(server, &rsp_iov, &oparms->fid->epoch, oparms->fid->lease_key, oplock, file_info, posix); + if (rc) + SMB2_close(xid, tcon, oparms->fid->persistent_fid, + oparms->fid->volatile_fid); trace_smb3_open_done(xid, rsp->PersistentFileId, tcon->tid, ses->Suid, oparms->create_options, oparms->desired_access, diff --git a/fs/smb/client/transport.c b/fs/smb/client/transport.c index e266859818a4..6e21b5f8754a 100644 --- a/fs/smb/client/transport.c +++ b/fs/smb/client/transport.c @@ -805,6 +805,18 @@ cifs_cancelled_callback(struct TCP_Server_Info *server, struct mid_q_entry *mid) release_mid(server, mid); } +static void +cifs_mark_compound_mids_cancelled(struct mid_q_entry **mid, int count) +{ + int i; + + for (i = 0; i < count; i++) { + spin_lock(&mid[i]->mid_lock); + mid[i]->wait_cancelled = true; + spin_unlock(&mid[i]->mid_lock); + } +} + /* * cifs_pick_channel - pick an eligible channel for network operations * @@ -865,6 +877,7 @@ compound_send_recv(const unsigned int xid, struct cifs_ses *ses, int *resp_buf_type, struct kvec *resp_iov) { int i, j, optype, rc = 0; + int num_processed = 0; struct mid_q_entry *mid[MAX_COMPOUND]; bool cancelled_mid[MAX_COMPOUND] = {false}; struct cifs_credits credits[MAX_COMPOUND] = { @@ -965,6 +978,10 @@ compound_send_recv(const unsigned int xid, struct cifs_ses *ses, if (rc < 0) { revert_current_mid(server, num_rqst); server->sequence_number -= 2; + for (i = 0; i < num_rqst; i++) { + delete_mid(server, mid[i]); + cancelled_mid[i] = true; + } } cifs_server_unlock(server); @@ -1011,6 +1028,14 @@ compound_send_recv(const unsigned int xid, struct cifs_ses *ses, break; } if (rc != 0) { + /* + * A completed CREATE earlier in the compound chain may have + * opened a remote handle even though a later wait was + * interrupted. Mark it cancelled so __release_mid() invokes + * the existing unmatched-open cleanup. + */ + cifs_mark_compound_mids_cancelled(mid, i); + for (; i < num_rqst; i++) { cifs_server_dbg(FYI, "Cancelling wait for mid %llu cmd: %d\n", mid[i]->mid, le16_to_cpu(mid[i]->command)); @@ -1033,6 +1058,14 @@ compound_send_recv(const unsigned int xid, struct cifs_ses *ses, rc = cifs_sync_mid_result(mid[i], server); if (rc != 0) { + /* + * A previous CREATE may have completed before this + * response failed. Mark it cancelled so its remote + * handle is closed when the mid is released. + */ + cifs_mark_compound_mids_cancelled(mid, i); + /* Keep their response buffers for cancelled-mid cleanup. */ + num_processed = 0; /* mark this mid as cancelled to not free it below */ cancelled_mid[i] = true; goto out; @@ -1042,13 +1075,24 @@ compound_send_recv(const unsigned int xid, struct cifs_ses *ses, mid[i]->mid_state != MID_RESPONSE_READY) { rc = smb_EIO1(smb_eio_trace_rx_mid_unready, mid[i]->mid_state); cifs_dbg(FYI, "Bad MID state?\n"); + cifs_mark_compound_mids_cancelled(mid, i); + num_processed = 0; goto out; } rc = server->ops->check_receive(mid[i], server, flags & CIFS_LOG_ERROR); + num_processed = i + 1; + } - if (resp_iov) { +out: + /* + * Delay moving response buffers out of their mids until response + * synchronization completes. This lets cancelled-mid cleanup inspect + * an earlier CREATE response if a later MID fails. + */ + if (resp_iov) { + for (i = 0; i < num_processed; i++) { buf = (char *)mid[i]->resp_buf; resp_iov[i].iov_base = buf; resp_iov[i].iov_len = mid[i]->resp_buf_size; @@ -1067,21 +1111,22 @@ compound_send_recv(const unsigned int xid, struct cifs_ses *ses, /* * Compounding is never used during session establish. */ - spin_lock(&ses->ses_lock); - if ((ses->ses_status == SES_NEW) || (optype & CIFS_NEG_OP) || (optype & CIFS_SESS_OP)) { - struct kvec iov = { - .iov_base = resp_iov[0].iov_base, - .iov_len = resp_iov[0].iov_len - }; - spin_unlock(&ses->ses_lock); - cifs_server_lock(server); - smb311_update_preauth_hash(ses, server, &iov, 1); - cifs_server_unlock(server); + if (num_processed == num_rqst) { spin_lock(&ses->ses_lock); + if ((ses->ses_status == SES_NEW) || (optype & CIFS_NEG_OP) || (optype & CIFS_SESS_OP)) { + struct kvec iov = { + .iov_base = resp_iov[0].iov_base, + .iov_len = resp_iov[0].iov_len + }; + spin_unlock(&ses->ses_lock); + cifs_server_lock(server); + smb311_update_preauth_hash(ses, server, &iov, 1); + cifs_server_unlock(server); + spin_lock(&ses->ses_lock); + } + spin_unlock(&ses->ses_lock); } - spin_unlock(&ses->ses_lock); -out: /* * This will dequeue all mids. After this it is important that the * demultiplex_thread will not process any of these mids any further. diff --git a/fs/squashfs/xz_wrapper.c b/fs/squashfs/xz_wrapper.c index 0a4ff3ec9c8c..6610af241449 100644 --- a/fs/squashfs/xz_wrapper.c +++ b/fs/squashfs/xz_wrapper.c @@ -57,10 +57,10 @@ static void *squashfs_xz_comp_opts(struct squashfs_sb_info *msblk, opts->dict_size = le32_to_cpu(comp_opts->dictionary_size); - /* the dictionary size should be 2^n or 2^n+2^(n+1) */ + /* the dictionary size should be positive and 2^n or 2^n+2^(n+1) */ n = ffs(opts->dict_size) - 1; - if (opts->dict_size != (1 << n) && opts->dict_size != (1 << n) + - (1 << (n + 1))) { + if (opts->dict_size <= 0 || (opts->dict_size != (1 << n) && + opts->dict_size != (1 << n) + (1 << (n + 1)))) { err = -EIO; goto out; } diff --git a/fs/super.c b/fs/super.c index a257b154f7f7..addf3ad54a58 100644 --- a/fs/super.c +++ b/fs/super.c @@ -103,7 +103,7 @@ static bool super_flags(const struct super_block *sb, unsigned int flags) * creation will succeed and SB_BORN is set by vfs_get_tree() or we're * woken and we'll see SB_DYING. * - * The caller must have acquired a temporary reference on @sb->s_count. + * The caller must have acquired a temporary reference on @sb->s_passive. * * Return: The function returns true if SB_BORN was set and with * s_umount held. The function returns false if SB_DYING was @@ -368,7 +368,7 @@ static struct super_block *alloc_super(struct file_system_type *type, int flags, spin_lock_init(&s->s_inode_wblist_lock); fserror_mount(s); - s->s_count = 1; + refcount_set(&s->s_passive, 1); atomic_set(&s->s_active, 1); mutex_init(&s->s_vfs_rename_mutex); lockdep_set_class(&s->s_vfs_rename_mutex, &type->s_vfs_rename_key); @@ -404,33 +404,28 @@ static struct super_block *alloc_super(struct file_system_type *type, int flags, /* Superblock refcounting */ /* - * Drop a superblock's refcount. The caller must hold sb_lock. + * Drop a superblock's passive reference. Must be called WITHOUT sb_lock held; + * put_super() acquires sb_lock itself when the final reference is dropped. */ -static void __put_super(struct super_block *s) +void put_super(struct super_block *s) { - if (!--s->s_count) { + if (refcount_dec_and_test(&s->s_passive)) { + struct file_system_type *type = s->s_type; + + spin_lock(&sb_lock); list_del_init(&s->s_list); + hlist_del_init(&s->s_instances); + spin_unlock(&sb_lock); + WARN_ON(s->s_dentry_lru.node); WARN_ON(s->s_inode_lru.node); WARN_ON(s->s_mounts); call_rcu(&s->rcu, destroy_super_rcu); + /* The unlink above may touch type->fs_supers, so drop it last. */ + put_filesystem(type); } } -/** - * put_super - drop a temporary reference to superblock - * @sb: superblock in question - * - * Drops a temporary reference, frees superblock if there's no - * references left. - */ -void put_super(struct super_block *sb) -{ - spin_lock(&sb_lock); - __put_super(sb); - spin_unlock(&sb_lock); -} - static void kill_super_notify(struct super_block *sb) { lockdep_assert_not_held(&sb->s_umount); @@ -439,24 +434,17 @@ static void kill_super_notify(struct super_block *sb) if (sb->s_flags & SB_DEAD) return; - /* - * Remove it from @fs_supers so it isn't found by new - * sget_fc() walkers anymore. Any concurrent mounter still - * managing to grab a temporary reference is guaranteed to - * already see SB_DYING and will wait until we notify them about - * SB_DEAD. - */ - spin_lock(&sb_lock); - hlist_del_init(&sb->s_instances); - spin_unlock(&sb_lock); - /* * Let concurrent mounts know that this thing is really dead. - * We don't need @sb->s_umount here as every concurrent caller - * will see SB_DYING and either discard the superblock or wait - * for SB_DEAD. + * sget_fc() skips SB_DEAD superblocks and calls test() under + * sb_lock, so set it under sb_lock: once we return no test() + * runs on this superblock anymore and none will start. Everyone + * else already saw SB_DYING and either discarded the superblock + * or waits for SB_DEAD. */ + spin_lock(&sb_lock); super_wake(sb, SB_DEAD); + spin_unlock(&sb_lock); } /** @@ -479,15 +467,10 @@ void deactivate_locked_super(struct super_block *s) kill_super_notify(s); - /* - * Since list_lru_destroy() may sleep, we cannot call it from - * put_super(), where we hold the sb_lock. Therefore we destroy - * the lru lists right now. - */ + /* list_lru_destroy() may sleep; put_super() callers may not. */ list_lru_destroy(&s->s_dentry_lru); list_lru_destroy(&s->s_inode_lru); - put_filesystem(fs); put_super(s); } else { super_unlock_excl(s); @@ -530,7 +513,7 @@ static bool grab_super(struct super_block *sb) { bool locked; - sb->s_count++; + refcount_inc(&sb->s_passive); spin_unlock(&sb_lock); locked = super_lock_excl(sb); if (locked) { @@ -557,7 +540,7 @@ static bool grab_super(struct super_block *sb) * lock held in read mode in case of success. On successful return, * the caller must drop the s_umount lock when done. * - * Note that unlike get_super() et.al. this one does *not* bump ->s_count. + * Note that unlike get_super() et.al. this one does *not* bump ->s_passive. * The reason why it's safe is that we are OK with doing trylock instead * of down_read(). There's a couple of places that are OK with that, but * it's very much not a general-purpose interface. @@ -674,12 +657,12 @@ void generic_shutdown_super(struct super_block *sb) } /* * Broadcast to everyone that grabbed a temporary reference to this - * superblock before we removed it from @fs_supers that the superblock - * is dying. Every walker of @fs_supers outside of sget_fc() will now - * discard this superblock and treat it as dead. + * superblock that it is dying. Every walker of @fs_supers outside + * of sget_fc() will now discard this superblock and treat it as + * dead. * - * We leave the superblock on @fs_supers so it can be found by - * sget_fc() until we passed sb->kill_sb(). + * sget_fc() keeps finding the superblock until SB_DEAD is set, so + * a concurrent mounter waits until we passed sb->kill_sb(). */ super_wake(sb, SB_DYING); super_unlock_excl(sb); @@ -758,6 +741,9 @@ struct super_block *sget_fc(struct fs_context *fc, spin_lock(&sb_lock); if (test) { hlist_for_each_entry(old, &fc->fs_type->fs_supers, s_instances) { + /* Only unlinked at the last passive reference. */ + if (super_flags(old, SB_DEAD)) + continue; if (test(old, fc)) goto share_extant_sb; } @@ -852,14 +838,17 @@ static void __iterate_supers(void (*f)(struct super_block *, void *), void *arg, struct super_block *sb, *p = NULL; bool excl = flags & SUPER_ITER_EXCL; - guard(spinlock)(&sb_lock); + spin_lock(&sb_lock); for (sb = first_super(flags); !list_entry_is_head(sb, &super_blocks, s_list); sb = next_super(sb, flags)) { if (super_flags(sb, SB_DYING)) continue; - sb->s_count++; + + if (!refcount_inc_not_zero(&sb->s_passive)) + continue; + spin_unlock(&sb_lock); if (flags & SUPER_ITER_UNLOCKED) { @@ -869,13 +858,14 @@ static void __iterate_supers(void (*f)(struct super_block *, void *), void *arg, super_unlock(sb, excl); } - spin_lock(&sb_lock); if (p) - __put_super(p); + put_super(p); p = sb; + spin_lock(&sb_lock); } + spin_unlock(&sb_lock); if (p) - __put_super(p); + put_super(p); } void iterate_supers(void (*f)(struct super_block *, void *), void *arg) @@ -904,7 +894,9 @@ void iterate_supers_type(struct file_system_type *type, if (super_flags(sb, SB_DYING)) continue; - sb->s_count++; + if (!refcount_inc_not_zero(&sb->s_passive)) + continue; + spin_unlock(&sb_lock); locked = super_lock_shared(sb); @@ -913,14 +905,14 @@ void iterate_supers_type(struct file_system_type *type, super_unlock_shared(sb); } - spin_lock(&sb_lock); if (p) - __put_super(p); + put_super(p); p = sb; + spin_lock(&sb_lock); } - if (p) - __put_super(p); spin_unlock(&sb_lock); + if (p) + put_super(p); } EXPORT_SYMBOL(iterate_supers_type); @@ -936,15 +928,17 @@ struct super_block *user_get_super(dev_t dev, bool excl) if (sb->s_dev != dev) continue; - sb->s_count++; + if (!refcount_inc_not_zero(&sb->s_passive)) + continue; + spin_unlock(&sb_lock); locked = super_lock(sb, excl); if (locked) return sb; + put_super(sb); spin_lock(&sb_lock); - __put_super(sb); break; } spin_unlock(&sb_lock); @@ -1375,9 +1369,7 @@ static struct super_block *bdev_super_lock(struct block_device *bdev, bool excl) lockdep_assert_not_held(&bdev->bd_disk->open_mutex); /* Make sure sb doesn't go away from under us */ - spin_lock(&sb_lock); - sb->s_count++; - spin_unlock(&sb_lock); + refcount_inc(&sb->s_passive); mutex_unlock(&bdev->bd_holder_lock); diff --git a/fs/xfs/libxfs/xfs_metafile.c b/fs/xfs/libxfs/xfs_metafile.c index 71f004e9dc64..1f54d39003c2 100644 --- a/fs/xfs/libxfs/xfs_metafile.c +++ b/fs/xfs/libxfs/xfs_metafile.c @@ -297,14 +297,14 @@ xfs_metafile_resv_init( goto out_unlock; /* - * Space taken by the per-AG metadata btrees are accounted on-disk as - * used space. We therefore only hide the space that is reserved but - * not used by the trees. + * Space taken by metadata btrees are accounted on-disk as used space. + * We therefore only hide the space that is reserved but not used by + * the trees. */ if (used > target) target = used; else if (target > dblocks_avail) - target = dblocks_avail; + target = max(dblocks_avail, used); hidden_space = target - used; error = xfs_dec_fdblocks(mp, hidden_space, true); diff --git a/fs/xfs/libxfs/xfs_rtrefcount_btree.c b/fs/xfs/libxfs/xfs_rtrefcount_btree.c index e2950dbe2068..dcc89b8e149b 100644 --- a/fs/xfs/libxfs/xfs_rtrefcount_btree.c +++ b/fs/xfs/libxfs/xfs_rtrefcount_btree.c @@ -617,7 +617,7 @@ xfs_rtrefcountbt_from_disk( fpp = xfs_rtrefcount_droot_ptr_addr(dblock, 1, maxrecs); tpp = xfs_rtrefcount_broot_ptr_addr(mp, rblock, 1, rblocklen); numrecs = be16_to_cpu(dblock->bb_numrecs); - memcpy(tkp, fkp, 2 * sizeof(*fkp) * numrecs); + memcpy(tkp, fkp, sizeof(*fkp) * numrecs); memcpy(tpp, fpp, sizeof(*fpp) * numrecs); } else { frp = xfs_rtrefcount_droot_rec_addr(dblock, 1); @@ -703,7 +703,7 @@ xfs_rtrefcountbt_to_disk( fpp = xfs_rtrefcount_broot_ptr_addr(mp, rblock, 1, rblocklen); tpp = xfs_rtrefcount_droot_ptr_addr(dblock, 1, maxrecs); numrecs = be16_to_cpu(rblock->bb_numrecs); - memcpy(tkp, fkp, 2 * sizeof(*fkp) * numrecs); + memcpy(tkp, fkp, sizeof(*fkp) * numrecs); memcpy(tpp, fpp, sizeof(*fpp) * numrecs); } else { frp = xfs_rtrefcount_rec_addr(rblock, 1); diff --git a/fs/xfs/scrub/dir.c b/fs/xfs/scrub/dir.c index 2a037aae904d..19d974c7e2b7 100644 --- a/fs/xfs/scrub/dir.c +++ b/fs/xfs/scrub/dir.c @@ -492,7 +492,7 @@ xchk_directory_data_bestfree( goto out; xchk_buffer_recheck(sc, bp); - if (xfs_has_crc(sc->mp)) { + if (!is_block && xfs_has_crc(sc->mp)) { struct xfs_dir3_data_hdr *hdr3 = bp->b_addr; if (hdr3->pad) diff --git a/fs/xfs/scrub/health.c b/fs/xfs/scrub/health.c index 2171bcf0f6c1..487ecc5f9f3c 100644 --- a/fs/xfs/scrub/health.c +++ b/fs/xfs/scrub/health.c @@ -202,9 +202,9 @@ xchk_update_health( * there's no sick flag defined for it, so we branch here ahead of the * mask check. */ - if (sc->sm->sm_type == XFS_SCRUB_TYPE_HEALTHY && - !(sc->sm->sm_flags & XFS_SCRUB_OFLAG_CORRUPT)) { - xchk_mark_all_healthy(sc->mp); + if (sc->sm->sm_type == XFS_SCRUB_TYPE_HEALTHY) { + if (!(sc->sm->sm_flags & XFS_SCRUB_OFLAG_CORRUPT)) + xchk_mark_all_healthy(sc->mp); return; } diff --git a/fs/xfs/scrub/inode.c b/fs/xfs/scrub/inode.c index 65b13e311916..46e9bf4a4317 100644 --- a/fs/xfs/scrub/inode.c +++ b/fs/xfs/scrub/inode.c @@ -607,7 +607,7 @@ xchk_dinode( } /* di_forkoff */ - if (XFS_DFORK_BOFF(dip) >= mp->m_sb.sb_inodesize) + if (dip->di_forkoff >= (XFS_LITINO(mp) >> 3)) xchk_ino_set_corrupt(sc, ino); if (naextents != 0 && dip->di_forkoff == 0) xchk_ino_set_corrupt(sc, ino); diff --git a/fs/xfs/scrub/inode_repair.c b/fs/xfs/scrub/inode_repair.c index 8bc508336aa5..b87c22146233 100644 --- a/fs/xfs/scrub/inode_repair.c +++ b/fs/xfs/scrub/inode_repair.c @@ -1702,7 +1702,7 @@ xrep_inode_blockcounts( &acount); if (error) return error; - if (count >= sc->mp->m_sb.sb_dblocks) + if (acount >= sc->mp->m_sb.sb_dblocks) return -EFSCORRUPTED; error = xrep_ino_ensure_extent_count(sc, XFS_ATTR_FORK, nextents); diff --git a/fs/xfs/scrub/newbt.c b/fs/xfs/scrub/newbt.c index c82f4631fd9c..584076b2a6ee 100644 --- a/fs/xfs/scrub/newbt.c +++ b/fs/xfs/scrub/newbt.c @@ -193,9 +193,11 @@ xrep_newbt_add_blocks( struct xrep_newbt_resv *resv; int error; - resv = kmalloc_obj(struct xrep_newbt_resv, XCHK_GFP_FLAGS); - if (!resv) - return -ENOMEM; + /* + * We have no way to clean up the allocated space *and* return an + * ENOMEM if we fail to allocate this control structure. + */ + resv = kmalloc_obj(struct xrep_newbt_resv, GFP_KERNEL | __GFP_NOFAIL); INIT_LIST_HEAD(&resv->list); resv->agbno = XFS_FSB_TO_AGBNO(mp, args->fsbno); diff --git a/fs/xfs/scrub/orphanage.c b/fs/xfs/scrub/orphanage.c index 3aca66869b80..21e31eeaa042 100644 --- a/fs/xfs/scrub/orphanage.c +++ b/fs/xfs/scrub/orphanage.c @@ -192,12 +192,16 @@ xrep_orphanage_create( /* Make sure the orphanage is owned by root. */ error = xrep_chown_orphanage(sc, XFS_I(orphanage_inode)); if (error) - goto out_dput_orphanage; + goto out_rele_orphanage; /* Stash the reference for later and bail out. */ sc->orphanage = XFS_I(orphanage_inode); sc->orphanage_ilock_flags = 0; + orphanage_inode = NULL; +out_rele_orphanage: + if (orphanage_inode) + xchk_irele(sc, XFS_I(orphanage_inode)); out_dput_orphanage: end_creating(orphanage_dentry); out_dput_root: diff --git a/fs/xfs/scrub/repair.c b/fs/xfs/scrub/repair.c index 11697a8b2a1d..c2a437416227 100644 --- a/fs/xfs/scrub/repair.c +++ b/fs/xfs/scrub/repair.c @@ -399,6 +399,7 @@ xrep_calc_rtgroup_resblks( struct xfs_mount *mp = sc->mp; struct xfs_scrub_metadata *sm = sc->sm; uint64_t usedlen; + xfs_extlen_t refcbt_sz = 0; xfs_extlen_t rmapbt_sz = 0; if (!(sm->sm_flags & XFS_SCRUB_IFLAG_REPAIR)) @@ -411,13 +412,27 @@ xrep_calc_rtgroup_resblks( usedlen = xfs_rtbxlen_to_blen(mp, xfs_rtgroup_extents(mp, sm->sm_agno)); ASSERT(usedlen <= XFS_MAX_RGBLOCKS); + if (xfs_has_reflink(mp)) + refcbt_sz = xfs_rtrefcountbt_calc_size(mp, usedlen); + if (xfs_has_rmapbt(mp)) rmapbt_sz = xfs_rtrmapbt_calc_size(mp, usedlen); + /* + * Guess how many blocks we need to rebuild the rmapbt. For + * non-reflink filesystems we can't have more records than used blocks. + * However, with reflink it's possible to have more than one rmap + * record per rtgroup block. We don't know how many rmaps there could + * be in the rtgroup, so we start off with what we hope is an generous + * over-estimation. + */ + if (refcbt_sz > 0 && rmapbt_sz > 0) + rmapbt_sz *= 2; + trace_xrep_calc_rtgroup_resblks_btsize(mp, sm->sm_agno, usedlen, - rmapbt_sz); + rmapbt_sz, refcbt_sz); - return rmapbt_sz; + return max(rmapbt_sz, refcbt_sz); } #endif /* CONFIG_XFS_RT */ diff --git a/fs/xfs/scrub/scrub.h b/fs/xfs/scrub/scrub.h index 6d7d3523b71f..7ade1d005efe 100644 --- a/fs/xfs/scrub/scrub.h +++ b/fs/xfs/scrub/scrub.h @@ -40,7 +40,7 @@ static inline int xchk_maybe_relax(struct xchk_relax *widget) return 0; widget->resched_nr = 0; - if (unlikely(widget->next_resched <= jiffies)) { + if (unlikely(time_after_eq(jiffies, widget->next_resched))) { cond_resched(); widget->next_resched = XCHK_RELAX_NEXT; } diff --git a/fs/xfs/scrub/trace.h b/fs/xfs/scrub/trace.h index 0f5adc293962..cb85f75ce101 100644 --- a/fs/xfs/scrub/trace.h +++ b/fs/xfs/scrub/trace.h @@ -2376,25 +2376,29 @@ TRACE_EVENT(xrep_calc_ag_resblks_btsize, #ifdef CONFIG_XFS_RT TRACE_EVENT(xrep_calc_rtgroup_resblks_btsize, TP_PROTO(struct xfs_mount *mp, xfs_rgnumber_t rgno, - xfs_rgblock_t usedlen, xfs_rgblock_t rmapbt_sz), - TP_ARGS(mp, rgno, usedlen, rmapbt_sz), + xfs_rgblock_t usedlen, xfs_rgblock_t rmapbt_sz, + xfs_rgblock_t refcbt_sz), + TP_ARGS(mp, rgno, usedlen, rmapbt_sz, refcbt_sz), TP_STRUCT__entry( __field(dev_t, dev) __field(xfs_rgnumber_t, rgno) __field(xfs_rgblock_t, usedlen) __field(xfs_rgblock_t, rmapbt_sz) + __field(xfs_rgblock_t, refcbt_sz) ), TP_fast_assign( __entry->dev = mp->m_super->s_dev; __entry->rgno = rgno; __entry->usedlen = usedlen; __entry->rmapbt_sz = rmapbt_sz; + __entry->refcbt_sz = refcbt_sz; ), - TP_printk("dev %d:%d rgno 0x%x usedlen %u rmapbt %u", + TP_printk("dev %d:%d rgno 0x%x usedlen %u rmapbt %u refcountbt %u", MAJOR(__entry->dev), MINOR(__entry->dev), __entry->rgno, __entry->usedlen, - __entry->rmapbt_sz) + __entry->rmapbt_sz, + __entry->refcbt_sz) ); #endif /* CONFIG_XFS_RT */ diff --git a/fs/xfs/xfs_dquot.c b/fs/xfs/xfs_dquot.c index b4f6c594808c..e696ee36c2e8 100644 --- a/fs/xfs/xfs_dquot.c +++ b/fs/xfs/xfs_dquot.c @@ -139,10 +139,14 @@ xfs_qm_adjust_dqlimits( dq->q_ino.softlimit = defq->ino.soft; if (!dq->q_ino.hardlimit) dq->q_ino.hardlimit = defq->ino.hard; - if (!dq->q_rtb.softlimit) + if (!dq->q_rtb.softlimit) { dq->q_rtb.softlimit = defq->rtb.soft; - if (!dq->q_rtb.hardlimit) + prealloc = 1; + } + if (!dq->q_rtb.hardlimit) { dq->q_rtb.hardlimit = defq->rtb.hard; + prealloc = 1; + } if (prealloc) xfs_dquot_set_prealloc_limits(dq); diff --git a/fs/xfs/xfs_exchrange.c b/fs/xfs/xfs_exchrange.c index c69ecd6a19de..fafb4e3f065c 100644 --- a/fs/xfs/xfs_exchrange.c +++ b/fs/xfs/xfs_exchrange.c @@ -633,6 +633,9 @@ xfs_exchrange_prep( if (error) return error; + if (fxr->flags & XFS_EXCHANGE_RANGE_DRY_RUN) + return 0; + trace_xfs_exchrange_flush(fxr, ip1, ip2); /* Flush the relevant ranges of both files. */ @@ -709,9 +712,11 @@ xfs_exchrange_contents( * other file write would do. This may involve turning on support for * logged xattrs if either file has security capabilities. */ - error = xfs_exchange_range_finish(fxr); - if (error) - goto out_unlock; + if (!(fxr->flags & XFS_EXCHANGE_RANGE_DRY_RUN)) { + error = xfs_exchange_range_finish(fxr); + if (error) + goto out_unlock; + } out_unlock: xfs_iunlock2_io_mmap(ip1, ip2); @@ -902,7 +907,7 @@ xfs_ioc_commit_range( if (copy_from_user(&args, argp, sizeof(args))) return -EFAULT; - if (args.flags & ~XFS_EXCHANGE_RANGE_ALL_FLAGS) + if (args.pad || (args.flags & ~XFS_EXCHANGE_RANGE_ALL_FLAGS)) return -EINVAL; if (kern_f->magic != XCR_FRESH_MAGIC) return -EBUSY; diff --git a/fs/xfs/xfs_healthmon.c b/fs/xfs/xfs_healthmon.c index acb85a1249a8..451575dceae0 100644 --- a/fs/xfs/xfs_healthmon.c +++ b/fs/xfs/xfs_healthmon.c @@ -245,7 +245,9 @@ xfs_healthmon_merge_events( case XFS_HEALTHMON_DIOWRITE: case XFS_HEALTHMON_DATALOST: /* logically adjacent file ranges can merge */ - if (existing->fino != new->fino || existing->fgen != new->fgen) + if (existing->fino != new->fino || + existing->fgen != new->fgen || + existing->error != new->error) return false; if (existing->fpos + existing->flen == new->fpos) { diff --git a/fs/xfs/xfs_icache.c b/fs/xfs/xfs_icache.c index 9d8dd30bd927..d462e17d4793 100644 --- a/fs/xfs/xfs_icache.c +++ b/fs/xfs/xfs_icache.c @@ -1656,7 +1656,7 @@ xfs_blockgc_free_dquots( do_work = true; } - if (XFS_IS_UQUOTA_ENFORCED(mp) && gdqp && xfs_dquot_lowsp(gdqp)) { + if (XFS_IS_GQUOTA_ENFORCED(mp) && gdqp && xfs_dquot_lowsp(gdqp)) { icw.icw_gid = make_kgid(mp->m_super->s_user_ns, gdqp->q_id); icw.icw_flags |= XFS_ICWALK_FLAG_GID; do_work = true; diff --git a/fs/xfs/xfs_log_recover.c b/fs/xfs/xfs_log_recover.c index fdb011e6ef60..8e4acebf72d6 100644 --- a/fs/xfs/xfs_log_recover.c +++ b/fs/xfs/xfs_log_recover.c @@ -2736,12 +2736,13 @@ xlog_recover_iunlink_bucket( { struct xfs_mount *mp = pag_mount(pag); struct xfs_inode *prev_ip = NULL; - struct xfs_inode *ip; xfs_agino_t prev_agino, agino; int error = 0; agino = be32_to_cpu(agi->agi_unlinked[bucket]); while (agino != NULLAGINO) { + struct xfs_inode *ip; + error = xfs_iget(mp, NULL, xfs_agino_to_ino(pag, agino), 0, 0, &ip); if (error) @@ -2750,11 +2751,11 @@ xlog_recover_iunlink_bucket( ASSERT(VFS_I(ip)->i_nlink == 0); ASSERT(VFS_I(ip)->i_mode != 0); xfs_iflags_clear(ip, XFS_IRECOVERY); - agino = ip->i_next_unlinked; if (prev_ip) { ip->i_prev_unlinked = prev_agino; xfs_irele(prev_ip); + prev_ip = NULL; /* * Ensure the inode is removed from the unlinked list @@ -2766,18 +2767,20 @@ xlog_recover_iunlink_bucket( * complete. */ error = xfs_inodegc_flush(mp); - if (error) - break; + if (error) { + xfs_irele(ip); + return error; + } } prev_agino = agino; + agino = ip->i_next_unlinked; prev_ip = ip; } if (prev_ip) { int error2; - ip->i_prev_unlinked = prev_agino; xfs_irele(prev_ip); error2 = xfs_inodegc_flush(mp); diff --git a/fs/xfs/xfs_qm.c b/fs/xfs/xfs_qm.c index 896b24f87ac9..f9c28d3864db 100644 --- a/fs/xfs/xfs_qm.c +++ b/fs/xfs/xfs_qm.c @@ -1432,16 +1432,22 @@ xfs_qm_flush_one( error = xfs_dquot_use_attached_buf(dqp, &bp); if (error) - goto out_unlock; + goto out_dqflock; if (!bp) { error = -EFSCORRUPTED; - goto out_unlock; + goto out_dqflock; } error = xfs_qm_dqflush(dqp, bp); if (!error) xfs_buf_delwri_queue(bp, buffer_list); xfs_buf_relse(bp); + mutex_unlock(&dqp->q_qlock); + xfs_qm_dqrele(dqp); + return error; + +out_dqflock: + xfs_dqfunlock(dqp); out_unlock: mutex_unlock(&dqp->q_qlock); xfs_qm_dqrele(dqp); diff --git a/fs/xfs/xfs_refcount_item.c b/fs/xfs/xfs_refcount_item.c index 8bccf89a7766..682c6e1b45e3 100644 --- a/fs/xfs/xfs_refcount_item.c +++ b/fs/xfs/xfs_refcount_item.c @@ -508,6 +508,7 @@ xfs_refcount_recover_work( struct xfs_cui_log_item *cuip = CUI_ITEM(lip); struct xfs_trans *tp; struct xfs_mount *mp = lip->li_log->l_mp; + unsigned int dblocks; bool isrt = xfs_cui_item_isrt(lip); int i; int error = 0; @@ -543,8 +544,11 @@ xfs_refcount_recover_work( * full btree split on either end of the refcount range. */ resv = xlog_recover_resv(&M_RES(mp)->tr_itruncate); - error = xfs_trans_alloc(mp, &resv, mp->m_refc_maxlevels * 2, 0, - XFS_TRANS_RESERVE, &tp); + if (isrt) + dblocks = mp->m_rtrefc_maxlevels * 2; + else + dblocks = mp->m_refc_maxlevels * 2; + error = xfs_trans_alloc(mp, &resv, dblocks, 0, XFS_TRANS_RESERVE, &tp); if (error) return error; diff --git a/fs/xfs/xfs_rmap_item.c b/fs/xfs/xfs_rmap_item.c index 2a3a73a8566d..000cff1ce324 100644 --- a/fs/xfs/xfs_rmap_item.c +++ b/fs/xfs/xfs_rmap_item.c @@ -573,6 +573,7 @@ xfs_rmap_recover_work( struct xfs_rui_log_item *ruip = RUI_ITEM(lip); struct xfs_trans *tp; struct xfs_mount *mp = lip->li_log->l_mp; + unsigned int dblocks; bool isrt = xfs_rui_item_isrt(lip); int i; int error = 0; @@ -596,8 +597,11 @@ xfs_rmap_recover_work( } resv = xlog_recover_resv(&M_RES(mp)->tr_itruncate); - error = xfs_trans_alloc(mp, &resv, mp->m_rmap_maxlevels, 0, - XFS_TRANS_RESERVE, &tp); + if (isrt) + dblocks = mp->m_rtrmap_maxlevels; + else + dblocks = mp->m_rmap_maxlevels; + error = xfs_trans_alloc(mp, &resv, dblocks, 0, XFS_TRANS_RESERVE, &tp); if (error) return error; diff --git a/include/linux/ethtool.h b/include/linux/ethtool.h index 12683b5d125e..97a1adbd9eae 100644 --- a/include/linux/ethtool.h +++ b/include/linux/ethtool.h @@ -944,6 +944,7 @@ struct kernel_ethtool_ts_info { #define ETHTOOL_OP_NEEDS_RTNL_SPAUSEPARAM BIT(6) #define ETHTOOL_OP_NEEDS_RTNL_RSS BIT(7) #define ETHTOOL_OP_NEEDS_RTNL_GLINK BIT(8) +#define ETHTOOL_OP_NEEDS_RTNL_TEST BIT(9) /** * struct ethtool_ops - optional netdev operations @@ -981,6 +982,7 @@ struct kernel_ethtool_ts_info { * - netdev_update_features() * - netif_set_real_num_tx_queues() * - ethtool_op_get_link() (syncs link watch under rtnl_lock) + * - netif_open() / netif_close() (used by @self_test) * * @get_drvinfo: Report driver/device information. Modern drivers no * longer have to implement this callback. Most fields are diff --git a/include/linux/fs/super_types.h b/include/linux/fs/super_types.h index ef7941e9dc79..68747182abf9 100644 --- a/include/linux/fs/super_types.h +++ b/include/linux/fs/super_types.h @@ -145,7 +145,7 @@ struct super_block { unsigned long s_magic; struct dentry *s_root; struct rw_semaphore s_umount; - int s_count; + refcount_t s_passive; atomic_t s_active; #ifdef CONFIG_SECURITY void *s_security; diff --git a/include/linux/if_vlan.h b/include/linux/if_vlan.h index 20cc16ea4e5a..4846032bf4ff 100644 --- a/include/linux/if_vlan.h +++ b/include/linux/if_vlan.h @@ -365,6 +365,9 @@ static inline int __vlan_insert_inner_tag(struct sk_buff *skb, const u8 meta_len = mac_len > ETH_TLEN ? skb_metadata_len(skb) : 0; struct vlan_ethhdr *veth; + if (unlikely(!pskb_may_pull(skb, mac_len))) + return -EINVAL; + if (skb_cow_head(skb, meta_len + VLAN_HLEN) < 0) return -ENOMEM; diff --git a/include/linux/mempool.h b/include/linux/mempool.h index a0fa6d43e0dc..6da502aef2f7 100644 --- a/include/linux/mempool.h +++ b/include/linux/mempool.h @@ -70,6 +70,13 @@ int mempool_alloc_bulk_noprof(struct mempool *pool, void **elem, #define mempool_alloc_bulk(...) \ alloc_hooks(mempool_alloc_bulk_noprof(__VA_ARGS__)) +/* + * Allocate a new element without dipping into the pool's reserves or + * waiting. Returns NULL on failure. + */ +#define mempool_alloc_noreserve(_pool, _gfp) \ + alloc_hooks((_pool)->alloc(_gfp, (_pool)->pool_data)) + void *mempool_alloc_preallocated(struct mempool *pool) __malloc; void mempool_free(void *element, struct mempool *pool); unsigned int mempool_free_bulk(struct mempool *pool, void **elem, diff --git a/include/linux/mm_types.h b/include/linux/mm_types.h index b18c2b2e7d2c..2b109cdaaad8 100644 --- a/include/linux/mm_types.h +++ b/include/linux/mm_types.h @@ -1211,7 +1211,7 @@ struct mm_struct { struct mm_mm_cid mm_cid; /* sched_cache related statistics */ - struct sched_cache_stat sc_stat; + struct sched_cache_group *sched_cache_grp; #ifdef CONFIG_MMU atomic_long_t pgtables_bytes; /* size of all page tables */ #endif @@ -1609,8 +1609,9 @@ static inline unsigned int mm_cid_size(void) #endif /* CONFIG_SCHED_MM_CID */ #ifdef CONFIG_SCHED_CACHE -void mm_init_sched(struct mm_struct *mm, - struct sched_cache_time __percpu *pcpu_sched); +int mm_init_sched(struct mm_struct *mm, + struct sched_cache_time __percpu *pcpu_sched); +void mm_destroy_sched(struct mm_struct *mm); static inline int mm_alloc_sched_noprof(struct mm_struct *mm) { @@ -1620,17 +1621,11 @@ static inline int mm_alloc_sched_noprof(struct mm_struct *mm) if (!pcpu_sched) return -ENOMEM; - mm_init_sched(mm, pcpu_sched); - return 0; + return mm_init_sched(mm, pcpu_sched); } #define mm_alloc_sched(...) alloc_hooks(mm_alloc_sched_noprof(__VA_ARGS__)) -static inline void mm_destroy_sched(struct mm_struct *mm) -{ - free_percpu(mm->sc_stat.pcpu_sched); - mm->sc_stat.pcpu_sched = NULL; -} #else /* !CONFIG_SCHED_CACHE */ static inline int mm_alloc_sched(struct mm_struct *mm) { return 0; } diff --git a/include/linux/perf_event.h b/include/linux/perf_event.h index 48d851fbd8ea..310681cccb50 100644 --- a/include/linux/perf_event.h +++ b/include/linux/perf_event.h @@ -1467,23 +1467,6 @@ static inline u32 perf_sample_data_size(struct perf_sample_data *data, return size; } -/* - * Clear all bitfields in the perf_branch_entry. - * The to and from fields are not cleared because they are - * systematically modified by caller. - */ -static inline void perf_clear_branch_entry_bitfields(struct perf_branch_entry *br) -{ - br->mispred = 0; - br->predicted = 0; - br->in_tx = 0; - br->abort = 0; - br->cycles = 0; - br->type = 0; - br->spec = PERF_BR_SPEC_NA; - br->reserved = 0; -} - extern void perf_output_sample(struct perf_output_handle *handle, struct perf_event_header *header, struct perf_sample_data *data, diff --git a/include/linux/sched.h b/include/linux/sched.h index 9341d647bbd1..1713fb6b33ed 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -2398,7 +2398,7 @@ struct sched_cache_time { unsigned long epoch; }; -struct sched_cache_stat { +struct sched_cache_group { struct sched_cache_time __percpu *pcpu_sched; raw_spinlock_t lock; unsigned long epoch; @@ -2406,11 +2406,13 @@ struct sched_cache_stat { unsigned long next_scan; unsigned long footprint; int cpu; + refcount_t refcnt; + struct rcu_head rcu; } ____cacheline_aligned_in_smp; #else -struct sched_cache_stat { }; +struct sched_cache_group { }; #endif diff --git a/include/linux/sched/topology.h b/include/linux/sched/topology.h index b5d9d7c2b8ad..f96812d71c51 100644 --- a/include/linux/sched/topology.h +++ b/include/linux/sched/topology.h @@ -281,9 +281,9 @@ static inline int task_node(const struct task_struct *p) } #ifdef CONFIG_SCHED_CACHE -extern void sched_update_llc_bytes(unsigned int cpu); +extern void sched_update_llc_bytes(const struct cpumask *cpus); #else -static inline void sched_update_llc_bytes(unsigned int cpu) { } +static inline void sched_update_llc_bytes(const struct cpumask *cpus) { } #endif #endif /* _LINUX_SCHED_TOPOLOGY_H */ diff --git a/include/linux/skbuff.h b/include/linux/skbuff.h index f7dd3db9459c..ce125cd24443 100644 --- a/include/linux/skbuff.h +++ b/include/linux/skbuff.h @@ -1834,22 +1834,6 @@ static inline void skb_zcopy_set(struct sk_buff *skb, struct ubuf_info *uarg, } } -static inline void skb_zcopy_set_nouarg(struct sk_buff *skb, void *val) -{ - skb_shinfo(skb)->destructor_arg = (void *)((uintptr_t) val | 0x1UL); - skb_shinfo(skb)->flags |= SKBFL_ZEROCOPY_FRAG; -} - -static inline bool skb_zcopy_is_nouarg(struct sk_buff *skb) -{ - return (uintptr_t) skb_shinfo(skb)->destructor_arg & 0x1UL; -} - -static inline void *skb_zcopy_get_nouarg(struct sk_buff *skb) -{ - return (void *)((uintptr_t) skb_shinfo(skb)->destructor_arg & ~0x1UL); -} - static inline void net_zcopy_put(struct ubuf_info *uarg) { if (uarg) @@ -1872,8 +1856,7 @@ static inline void skb_zcopy_clear(struct sk_buff *skb, bool zerocopy_success) struct ubuf_info *uarg = skb_zcopy(skb); if (uarg) { - if (!skb_zcopy_is_nouarg(skb)) - uarg->ops->complete(skb, uarg, zerocopy_success); + uarg->ops->complete(skb, uarg, zerocopy_success); skb_shinfo(skb)->flags &= ~SKBFL_ALL_ZEROCOPY; } @@ -4372,7 +4355,10 @@ skb_header_pointer_careful(const struct sk_buff *skb, int offset, static inline void * __must_check skb_pointer_if_linear(const struct sk_buff *skb, int offset, int len) { - if (likely(skb_headlen(skb) - offset >= len)) + unsigned int uoffset = (unsigned int)offset; + + if (likely(uoffset <= skb_headlen(skb) && + (unsigned int)len <= skb_headlen(skb) - uoffset)) return skb->data + offset; return NULL; } diff --git a/include/net/dst.h b/include/net/dst.h index 307073eae7f8..dbedfe72e1fd 100644 --- a/include/net/dst.h +++ b/include/net/dst.h @@ -455,7 +455,8 @@ static inline unsigned int dst_dev_overhead(struct dst_entry *dst, struct sk_buff *skb) { if (likely(dst)) - return LL_RESERVED_SPACE(dst->dev); + return max_t(unsigned int, skb->mac_len, + LL_RESERVED_SPACE(dst->dev)); return skb->mac_len; } diff --git a/include/net/gue.h b/include/net/gue.h index caefd6da8693..d377155fd0b3 100644 --- a/include/net/gue.h +++ b/include/net/gue.h @@ -84,8 +84,9 @@ static inline size_t guehdr_priv_flags_len(__be32 flags) } /* Validate standard and private flags. Returns non-zero (meaning invalid) - * if there is an unknown standard or private flags, or the options length for - * the flags exceeds the options length specific in hlen of the GUE header. + * if there is an unknown standard or private flags, if the options length for + * the flags exceeds the options length specified in hlen of the GUE header, or + * if a private option contains invalid data. */ static inline int validate_gue_flags(struct guehdr *guehdr, size_t optlen) { @@ -103,8 +104,8 @@ static inline int validate_gue_flags(struct guehdr *guehdr, size_t optlen) /* Private flags are last four bytes accounted in * guehdr_flags_len */ - __be32 pflags = *(__be32 *)((void *)&guehdr[1] + - len - GUE_LEN_PRIV); + void *data = (void *)&guehdr[1] + len; + __be32 pflags = *(__be32 *)(data - GUE_LEN_PRIV); if (pflags & ~GUE_PFLAGS_ALL) return 1; @@ -112,6 +113,16 @@ static inline int validate_gue_flags(struct guehdr *guehdr, size_t optlen) len += guehdr_priv_flags_len(pflags); if (len > optlen) return 1; + + if (pflags & GUE_PFLAG_REMCSUM) { + __be16 *pd = data; + + /* The field offset pd[1] must not be less + * than the start pd[0]. + */ + if (ntohs(pd[1]) < ntohs(pd[0])) + return 1; + } } return 0; diff --git a/include/net/ip6_route.h b/include/net/ip6_route.h index 745ab62154ac..0a6c4ceb95a7 100644 --- a/include/net/ip6_route.h +++ b/include/net/ip6_route.h @@ -101,12 +101,12 @@ static inline struct dst_entry *ip6_route_output(struct net *net, } /* Only conditionally release dst if flags indicates - * !RT6_LOOKUP_F_DST_NOREF or dst is in uncached_list. + * !RT6_LOOKUP_F_DST_NOREF or dst is uncached. */ static inline void ip6_rt_put_flags(struct rt6_info *rt, int flags) { if (!(flags & RT6_LOOKUP_F_DST_NOREF) || - !list_empty(&rt->dst.rt_uncached)) + rt->dst.rt_uncached_list) ip6_rt_put(rt); } diff --git a/include/net/netfilter/nf_conntrack.h b/include/net/netfilter/nf_conntrack.h index bc42dd0e10e6..c39425e54d87 100644 --- a/include/net/netfilter/nf_conntrack.h +++ b/include/net/netfilter/nf_conntrack.h @@ -185,6 +185,11 @@ static inline void nf_ct_put(struct nf_conn *ct) nf_ct_destroy(&ct->ct_general); } +static inline bool nf_ct_shared(const struct nf_conn *ct) +{ + return refcount_read(&ct->ct_general.use) > 1; +} + /* load module; enable/disable conntrack in this namespace */ int nf_ct_netns_get(struct net *net, u8 nfproto); void nf_ct_netns_put(struct net *net, u8 nfproto); diff --git a/include/net/nfc/hci.h b/include/net/nfc/hci.h index 756c11084f65..86ed63e5d533 100644 --- a/include/net/nfc/hci.h +++ b/include/net/nfc/hci.h @@ -144,7 +144,7 @@ struct nfc_hci_dev { data_exchange_cb_t async_cb; void *async_cb_context; - u8 *gb; + u8 gb[NFC_MAX_GT_LEN]; size_t gb_len; unsigned long quirks; diff --git a/include/net/nfc/nfc.h b/include/net/nfc/nfc.h index c54df042db6b..bcafab5c53e5 100644 --- a/include/net/nfc/nfc.h +++ b/include/net/nfc/nfc.h @@ -273,7 +273,8 @@ struct sk_buff *nfc_alloc_recv_skb(unsigned int size, gfp_t gfp); int nfc_set_remote_general_bytes(struct nfc_dev *dev, const u8 *gt, u8 gt_len); -u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, size_t *gb_len); +u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, u8 *out_gb, + size_t gb_max_len, size_t *gb_len); int nfc_fw_download_done(struct nfc_dev *dev, const char *firmware_name, u32 result); diff --git a/include/net/tcp.h b/include/net/tcp.h index 436495ff2271..4416cdf9bf30 100644 --- a/include/net/tcp.h +++ b/include/net/tcp.h @@ -1232,9 +1232,9 @@ static inline bool tcp_skb_can_collapse_to(const struct sk_buff *skb) static inline bool tcp_skb_can_collapse(const struct sk_buff *to, const struct sk_buff *from) { - /* skb_cmp_decrypted() not needed, use tcp_write_collapse_fence() */ return likely(tcp_skb_can_collapse_to(to) && mptcp_skb_can_collapse(to, from) && + !skb_cmp_decrypted(to, from) && skb_pure_zcopy_same(to, from) && skb_frags_readable(to) == skb_frags_readable(from)); } @@ -2327,7 +2327,7 @@ static inline void tcp_rtx_queue_unlink_and_free(struct sk_buff *skb, struct soc static inline void tcp_write_collapse_fence(struct sock *sk) { - struct sk_buff *skb = tcp_write_queue_tail(sk); + struct sk_buff *skb = tcp_write_queue_tail(sk) ?: tcp_rtx_queue_tail(sk); if (skb) TCP_SKB_CB(skb)->eor = 1; diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c index 022d5a594dc9..c4df013bf006 100644 --- a/kernel/bpf/btf.c +++ b/kernel/bpf/btf.c @@ -4269,13 +4269,10 @@ int btf_check_and_fixup_fields(const struct btf *btf, struct btf_record *rec) { int i; - /* There are three types that signify ownership of some other type: - * kptr_ref, bpf_list_head, bpf_rb_root. - * kptr_ref only supports storing kernel types, which can't store - * references to program allocated local types. - * - * Hence we only need to ensure that bpf_{list_head,rb_root} ownership - * does not form cycles. + /* + * Check fields which require the complete BTF and initialize runtime + * metadata. Ownership relationships are validated after every record has + * been fixed up. */ if (IS_ERR_OR_NULL(rec) || !(rec->field_mask & (BPF_GRAPH_ROOT | BPF_UPTR))) return 0; @@ -4306,51 +4303,88 @@ int btf_check_and_fixup_fields(const struct btf *btf, struct btf_record *rec) if (!meta) return -EFAULT; rec->fields[i].graph_root.value_rec = meta->record; + } + return 0; +} - /* We need to set value_rec for all root types, but no need - * to check ownership cycle for a type unless it's also a - * node type. - */ - if (!(rec->field_mask & BPF_GRAPH_NODE)) +static int btf_owned_type_idx(const struct btf *btf, struct btf_struct_metas *tab, + const struct btf_field *field) +{ + struct btf_struct_meta *meta; + u32 btf_id; + + if (field->type & BPF_GRAPH_ROOT) { + btf_id = field->graph_root.value_btf_id; + } else if (field->type == BPF_KPTR_REF || field->type == BPF_KPTR_PERCPU) { + if (btf_is_kernel(field->kptr.btf)) + return -ENOENT; + btf_id = field->kptr.btf_id; + } else { + return -ENOENT; + } + + meta = btf_find_struct_meta(btf, btf_id); + if (!meta) + return field->type & BPF_GRAPH_ROOT ? -EFAULT : -ENOENT; + return meta - tab->types; +} + +/* + * Each ownership edge adds kernel frames through bpf_obj_free_fields() and + * __bpf_obj_drop_impl(). Keep the bound deliberately small because object + * destruction can itself run below a BPF call chain. A final pointee without + * special fields is not present in the struct metadata table and adds only a + * non-recursing drop. + */ +#define BTF_MAX_OWNERSHIP_DEPTH 8 + +static int btf_ownership_depth(const struct btf *btf, + struct btf_struct_metas *tab, u8 *depth, + int idx, int depth_left) +{ + const struct btf_record *rec = tab->types[idx].record; + int i, ret, max_depth = 0; + + if (!depth_left) + return -ELOOP; + if (depth[idx]) + goto done; + + for (i = 0; i < rec->cnt; i++) { + ret = btf_owned_type_idx(btf, tab, &rec->fields[i]); + if (ret == -ENOENT) continue; + if (ret < 0) + return ret; + ret = btf_ownership_depth(btf, tab, depth, ret, depth_left - 1); + if (ret < 0) + return ret; + max_depth = max(max_depth, ret); + } + depth[idx] = max_depth + 1; +done: + return depth[idx] > depth_left ? -ELOOP : depth[idx]; +} - /* We need to ensure ownership acyclicity among all types. The - * proper way to do it would be to topologically sort all BTF - * IDs based on the ownership edges, since there can be multiple - * bpf_{list_head,rb_node} in a type. Instead, we use the - * following resaoning: - * - * - A type can only be owned by another type in user BTF if it - * has a bpf_{list,rb}_node. Let's call these node types. - * - A type can only _own_ another type in user BTF if it has a - * bpf_{list_head,rb_root}. Let's call these root types. - * - * We ensure that if a type is both a root and node, its - * element types cannot be root types. - * - * To ensure acyclicity: - * - * When A is an root type but not a node, its ownership - * chain can be: - * A -> B -> C - * Where: - * - A is an root, e.g. has bpf_rb_root. - * - B is both a root and node, e.g. has bpf_rb_node and - * bpf_list_head. - * - C is only an root, e.g. has bpf_list_node - * - * When A is both a root and node, some other type already - * owns it in the BTF domain, hence it can not own - * another root type through any of the ownership edges. - * A -> B - * Where: - * - A is both an root and node. - * - B is only an node. - */ - if (meta->record->field_mask & BPF_GRAPH_ROOT) - return -ELOOP; +static int btf_check_ownership_depth(const struct btf *btf, + struct btf_struct_metas *tab) +{ + u8 *depth; + int i, ret = 0; + + depth = kvcalloc(tab->cnt, sizeof(*depth), GFP_KERNEL | __GFP_NOWARN); + if (!depth) + return -ENOMEM; + + for (i = 0; i < tab->cnt; i++) { + ret = btf_ownership_depth(btf, tab, depth, i, + BTF_MAX_OWNERSHIP_DEPTH); + if (ret < 0) + break; + ret = 0; } - return 0; + kvfree(depth); + return ret; } static void __btf_struct_show(const struct btf *btf, const struct btf_type *t, @@ -6045,6 +6079,10 @@ static struct btf *btf_parse(const union bpf_attr *attr, bpfptr_t uattr, if (err < 0) goto errout_meta; } + + err = btf_check_ownership_depth(btf, struct_meta_tab); + if (err < 0) + goto errout_meta; } err = bpf_log_attr_finalize(attr_log, &env->log); @@ -7191,7 +7229,7 @@ static int btf_struct_walk(struct bpf_verifier_log *log, const struct btf *btf, if (btf_type_is_int(t)) return WALK_SCALAR; - if (!btf_type_is_struct(t)) + if (!btf_type_is_struct(t) || !t->size) goto error; off = (off - moff) % t->size; diff --git a/kernel/bpf/core.c b/kernel/bpf/core.c index 883cb7a800b1..71ee96a77ae2 100644 --- a/kernel/bpf/core.c +++ b/kernel/bpf/core.c @@ -19,6 +19,7 @@ #include #include +#include #include #include #include @@ -1612,6 +1613,8 @@ struct bpf_prog *bpf_jit_blind_constants(struct bpf_verifier_env *env, struct bp * fix it up here on error. */ bpf_jit_prog_release_other(prog, clone); + if (env && fatal_signal_pending(current)) + return ERR_PTR(-EINTR); return IS_ERR(tmp) ? tmp : ERR_PTR(-ENOMEM); } @@ -2641,11 +2644,14 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc orig_prog = prog; prog = bpf_jit_blind_constants(env, prog); /* - * If blinding was requested and we failed during blinding, we must fall - * back to the interpreter. + * Fall back to the interpreter after blinding failures, except when + * the loader was killed. */ - if (IS_ERR(prog)) + if (IS_ERR(prog)) { + if (PTR_ERR(prog) == -EINTR) + return prog; goto out_restore; + } prog = bpf_int_jit_compile(env, prog); if (prog->jited) { @@ -2668,6 +2674,8 @@ static struct bpf_prog *bpf_prog_jit_compile(struct bpf_verifier_env *env, struc struct bpf_prog *__bpf_prog_select_runtime(struct bpf_verifier_env *env, struct bpf_prog *fp, int *err) { + struct bpf_prog *jit_prog; + /* In case of BPF to BPF calls, verifier did all the prep * work with regards to JITing, etc. */ @@ -2690,7 +2698,12 @@ struct bpf_prog *__bpf_prog_select_runtime(struct bpf_verifier_env *env, struct if (*err) return fp; - fp = bpf_prog_jit_compile(env, fp); + jit_prog = bpf_prog_jit_compile(env, fp); + if (IS_ERR(jit_prog)) { + *err = PTR_ERR(jit_prog); + return fp; + } + fp = jit_prog; bpf_prog_jit_attempt_done(fp); if (!fp->jited && jit_needed) { *err = -ENOTSUPP; diff --git a/kernel/bpf/crypto.c b/kernel/bpf/crypto.c index 51f89cecefb4..3f3fe2450fc6 100644 --- a/kernel/bpf/crypto.c +++ b/kernel/bpf/crypto.c @@ -149,8 +149,9 @@ bpf_crypto_ctx_create(const struct bpf_crypto_params *params, u32 params__sz, const struct bpf_crypto_type *type; struct bpf_crypto_ctx *ctx; - if (!params || params->reserved[0] || params->reserved[1] || - params__sz != sizeof(struct bpf_crypto_params)) { + if (!params || + params__sz != sizeof(struct bpf_crypto_params) || + params->reserved[0] || params->reserved[1]) { *err = -EINVAL; return NULL; } diff --git a/kernel/bpf/fixups.c b/kernel/bpf/fixups.c index 48acb61854ed..19c1b8f1b317 100644 --- a/kernel/bpf/fixups.c +++ b/kernel/bpf/fixups.c @@ -8,6 +8,7 @@ #include #include #include +#include #include #include "disasm.h" @@ -240,12 +241,28 @@ static void adjust_poke_descs(struct bpf_prog *prog, u32 off, u32 len) } } +/* + * Some post-verification instruction rewriting passes require an + * O(prog->len) operation per instruction. Keep their shared primitives + * killable and preemptible. + */ +static bool bpf_rewrite_must_abort(void) +{ + if (fatal_signal_pending(current)) + return true; + cond_resched(); + return false; +} + struct bpf_prog *bpf_patch_insn_data(struct bpf_verifier_env *env, u32 off, const struct bpf_insn *patch, u32 len) { struct bpf_prog *new_prog; struct bpf_insn_aux_data *new_data = NULL; + if (bpf_rewrite_must_abort()) + return NULL; + if (len > 1) { new_data = vrealloc(env->insn_aux_data, array_size(env->prog->len + len - 1, @@ -457,6 +474,9 @@ static int verifier_remove_insns(struct bpf_verifier_env *env, u32 off, u32 cnt) unsigned int orig_prog_len = env->prog->len; int err; + if (bpf_rewrite_must_abort()) + return -EINTR; + if (bpf_prog_is_offloaded(env->prog->aux)) bpf_prog_offload_remove_insns(env, off, cnt); @@ -1316,7 +1336,7 @@ int bpf_jit_subprogs(struct bpf_verifier_env *env) } prog = bpf_jit_blind_constants(env, prog); if (IS_ERR(prog)) { - err = -ENOMEM; + err = PTR_ERR(prog); prog = orig_prog; goto out_restore; } @@ -1395,7 +1415,7 @@ int bpf_fixup_call_args(struct bpf_verifier_env *env) err = bpf_jit_subprogs(env); if (err == 0) return 0; - if (err == -EFAULT) + if (err == -EFAULT || err == -EINTR) return err; } #ifndef CONFIG_BPF_JIT_ALWAYS_ON diff --git a/kernel/bpf/hashtab.c b/kernel/bpf/hashtab.c index 59406da06424..61ef6191f79a 100644 --- a/kernel/bpf/hashtab.c +++ b/kernel/bpf/hashtab.c @@ -1055,14 +1055,17 @@ static void pcpu_init_value(struct bpf_htab *htab, void __percpu *pptr, /* When not setting the initial value on all cpus, zero-fill element * values for other cpus. Otherwise, bpf program has no way to ensure * known initial values for cpus other than current one - * (onallcpus=false always when coming from bpf prog). + * (onallcpus=false always when coming from bpf prog, + * map_flags & BPF_F_CPU when coming from syscall but setting + * only one cpu). */ - if (!onallcpus) { - int current_cpu = raw_smp_processor_id(); + if (!onallcpus || (map_flags & BPF_F_CPU)) { + int init_cpu = (map_flags & BPF_F_CPU) ? map_flags >> 32 : + raw_smp_processor_id(); int cpu; for_each_possible_cpu(cpu) { - if (cpu == current_cpu) + if (cpu == init_cpu) copy_map_value(&htab->map, per_cpu_ptr(pptr, cpu), value); else /* Since elem is preallocated, we cannot touch special fields */ zero_map_value(&htab->map, per_cpu_ptr(pptr, cpu)); @@ -1773,6 +1776,12 @@ static int htab_lru_percpu_map_lookup_and_delete_elem(struct bpf_map *map, flags); } +/* + * Max consecutive empty buckets to walk in one RCU + + * instrumentation-disabled section before rescheduling. + */ +#define HTAB_BATCH_EMPTY_RESCHED 64 + static int __htab_map_lookup_and_delete_batch(struct bpf_map *map, const union bpf_attr *attr, @@ -1794,6 +1803,7 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map, unsigned long flags = 0; bool locked = false; struct htab_elem *l; + u32 empty_cnt = 0; struct bucket *b; int ret = 0; @@ -1972,30 +1982,41 @@ __htab_map_lookup_and_delete_batch(struct bpf_map *map, } next_batch: - /* If we are not copying data, we can go to next bucket and avoid - * unlocking the rcu. + /* + * If we are not copying data, we can go to next bucket and avoid + * unlocking the rcu. Bound the walk though: after + * HTAB_BATCH_EMPTY_RESCHED consecutive empty buckets, fully exit + * the critical section (no locks are held here) and reschedule. */ if (!bucket_cnt && (batch + 1 < htab->n_buckets)) { batch++; - goto again_nocopy; + if (++empty_cnt < HTAB_BATCH_EMPTY_RESCHED) + goto again_nocopy; + empty_cnt = 0; + rcu_read_unlock(); + bpf_enable_instrumentation(); + cond_resched_tasks_rcu_qs(); + goto again; } rcu_read_unlock(); bpf_enable_instrumentation(); - if (bucket_cnt && (copy_to_user(ukeys + total * key_size, keys, - key_size * bucket_cnt) || - copy_to_user(uvalues + total * value_size, values, - value_size * bucket_cnt))) { + if (bucket_cnt && (copy_to_user(ukeys + (size_t)total * key_size, keys, + (size_t)key_size * bucket_cnt) || + copy_to_user(uvalues + (size_t)total * value_size, values, + (size_t)value_size * bucket_cnt))) { ret = -EFAULT; goto after_loop; } total += bucket_cnt; + empty_cnt = 0; batch++; if (batch >= htab->n_buckets) { ret = -ENOENT; goto after_loop; } + cond_resched_tasks_rcu_qs(); goto again; after_loop: diff --git a/kernel/bpf/memalloc.c b/kernel/bpf/memalloc.c index e9662db7198f..8a8f088e83e6 100644 --- a/kernel/bpf/memalloc.c +++ b/kernel/bpf/memalloc.c @@ -119,6 +119,7 @@ struct bpf_mem_cache { struct llist_head waiting_for_gp_ttrace; struct rcu_head rcu_ttrace; atomic_t call_rcu_ttrace_in_progress; + raw_spinlock_t lock; }; struct bpf_mem_caches { @@ -214,25 +215,24 @@ static void alloc_bulk(struct bpf_mem_cache *c, int cnt, int node, bool atomic) gfp = __GFP_NOWARN | __GFP_ACCOUNT; gfp |= atomic ? GFP_NOWAIT : GFP_KERNEL; - for (i = 0; i < cnt; i++) { - /* - * For every 'c' llist_del_first(&c->free_by_rcu_ttrace); is - * done only by one CPU == current CPU. Other CPUs might - * llist_add() and llist_del_all() in parallel. - */ - obj = llist_del_first(&c->free_by_rcu_ttrace); - if (!obj) - break; - add_obj_to_free_list(c, obj); - } - if (i >= cnt) - return; + /* + * c->lock serializes concurrent llist_del_first() against + * llist_del_all() in __free_rcu() and do_call_rcu_ttrace(). + */ + scoped_guard(raw_spinlock_irqsave, &c->lock) { + for (i = 0; i < cnt; i++) { + obj = llist_del_first(&c->free_by_rcu_ttrace); + if (!obj) + break; + add_obj_to_free_list(c, obj); + } - for (; i < cnt; i++) { - obj = llist_del_first(&c->waiting_for_gp_ttrace); - if (!obj) - break; - add_obj_to_free_list(c, obj); + for (; i < cnt; i++) { + obj = llist_del_first(&c->waiting_for_gp_ttrace); + if (!obj) + break; + add_obj_to_free_list(c, obj); + } } if (i >= cnt) return; @@ -279,8 +279,12 @@ static int free_all(struct bpf_mem_cache *c, struct llist_node *llnode, bool per static void __free_rcu(struct rcu_head *head) { struct bpf_mem_cache *c = container_of(head, struct bpf_mem_cache, rcu_ttrace); + struct llist_node *llnode; + + scoped_guard(raw_spinlock_irqsave, &c->lock) + llnode = llist_del_all(&c->waiting_for_gp_ttrace); - free_all(c, llist_del_all(&c->waiting_for_gp_ttrace), !!c->percpu_size); + free_all(c, llnode, !!c->percpu_size); atomic_set(&c->call_rcu_ttrace_in_progress, 0); } @@ -300,7 +304,8 @@ static void do_call_rcu_ttrace(struct bpf_mem_cache *c) if (atomic_xchg(&c->call_rcu_ttrace_in_progress, 1)) { if (unlikely(READ_ONCE(c->draining))) { - llnode = llist_del_all(&c->free_by_rcu_ttrace); + scoped_guard(raw_spinlock_irqsave, &c->lock) + llnode = llist_del_all(&c->free_by_rcu_ttrace); free_all(c, llnode, !!c->percpu_size); } return; @@ -535,6 +540,7 @@ int bpf_mem_alloc_init(struct bpf_mem_alloc *ma, int size, bool percpu) c->objcg = objcg; c->percpu_size = percpu_size; c->tgt = c; + raw_spin_lock_init(&c->lock); init_refill_work(c); prefill_mem_cache(c, cpu); } @@ -557,7 +563,7 @@ int bpf_mem_alloc_init(struct bpf_mem_alloc *ma, int size, bool percpu) c->objcg = objcg; c->percpu_size = percpu_size; c->tgt = c; - + raw_spin_lock_init(&c->lock); init_refill_work(c); prefill_mem_cache(c, cpu); } @@ -609,7 +615,7 @@ int bpf_mem_alloc_percpu_unit_init(struct bpf_mem_alloc *ma, int size) c->objcg = objcg; c->percpu_size = percpu_size; c->tgt = c; - + raw_spin_lock_init(&c->lock); init_refill_work(c); prefill_mem_cache(c, cpu); } diff --git a/kernel/bpf/offload.c b/kernel/bpf/offload.c index 0d6f5569588c..d855399812ee 100644 --- a/kernel/bpf/offload.c +++ b/kernel/bpf/offload.c @@ -698,6 +698,8 @@ static bool __bpf_offload_dev_match(struct bpf_prog *prog, return false; if (offload->netdev == netdev) return true; + if (!bpf_prog_is_offloaded(prog->aux)) + return false; ondev1 = bpf_offload_find_netdev(offload->netdev); ondev2 = bpf_offload_find_netdev(netdev); diff --git a/kernel/bpf/states.c b/kernel/bpf/states.c index 66fb11b6c6a7..6c88ad95b63b 100644 --- a/kernel/bpf/states.c +++ b/kernel/bpf/states.c @@ -635,6 +635,9 @@ static bool regsafe(struct bpf_verifier_env *env, struct bpf_reg_state *rold, /* id relations must be preserved */ if (!check_ids(rold->id, rcur->id, idmap)) return false; + /* Preserve displacements between pointers sharing an ID. */ + if (rold->id && rold->r64.base != rcur->r64.base) + return false; /* new val must satisfy old val knowledge */ return range_within(rold, rcur) && tnum_in(rold->var_off, rcur->var_off); diff --git a/kernel/bpf/syscall.c b/kernel/bpf/syscall.c index c7cb336fb064..77b3252270f3 100644 --- a/kernel/bpf/syscall.c +++ b/kernel/bpf/syscall.c @@ -2033,7 +2033,7 @@ int generic_map_delete_batch(struct bpf_map *map, for (cp = 0; cp < max_count; cp++) { err = -EFAULT; - if (copy_from_user(key, keys + cp * map->key_size, + if (copy_from_user(key, keys + (size_t)cp * map->key_size, map->key_size)) break; @@ -2095,9 +2095,9 @@ int generic_map_update_batch(struct bpf_map *map, struct file *map_file, for (cp = 0; cp < max_count; cp++) { err = -EFAULT; - if (copy_from_user(key, keys + cp * map->key_size, + if (copy_from_user(key, keys + (size_t)cp * map->key_size, map->key_size) || - copy_from_user(value, values + cp * value_size, value_size)) + copy_from_user(value, values + (size_t)cp * value_size, value_size)) break; err = bpf_map_update_value(map, map_file, key, value, @@ -2176,12 +2176,12 @@ int generic_map_lookup_batch(struct bpf_map *map, if (err) goto free_buf; - if (copy_to_user(keys + cp * map->key_size, key, + if (copy_to_user(keys + (size_t)cp * map->key_size, key, map->key_size)) { err = -EFAULT; goto free_buf; } - if (copy_to_user(values + cp * value_size, value, value_size)) { + if (copy_to_user(values + (size_t)cp * value_size, value, value_size)) { err = -EFAULT; goto free_buf; } @@ -6106,7 +6106,10 @@ struct bpf_link *bpf_link_get_curr_or_next(u32 *id) again: link = idr_get_next(&link_idr, id); if (link) { - link = bpf_link_inc_not_zero(link); + if (link->id) + link = bpf_link_inc_not_zero(link); + else + link = ERR_PTR(-EAGAIN); if (IS_ERR(link)) { (*id)++; goto again; diff --git a/kernel/bpf/verifier.c b/kernel/bpf/verifier.c index 1a0cd37b03cd..ccf735d74dc8 100644 --- a/kernel/bpf/verifier.c +++ b/kernel/bpf/verifier.c @@ -582,7 +582,7 @@ static int stack_slot_obj_get_spi(struct bpf_verifier_env *env, struct bpf_reg_s } off = reg->var_off.value; - if (off % BPF_REG_SIZE) { + if (off >= 0 || off % BPF_REG_SIZE) { verbose(env, "cannot pass in %s at an offset=%d\n", obj_kind, off); return -EINVAL; } @@ -2876,6 +2876,8 @@ static int check_subprogs(struct bpf_verifier_env *env) subprog[cur_subprog].exit_idx = i; goto next; } + if (insn_is_gotox(&insn[i])) + goto next; off = i + bpf_jmp_offset(&insn[i]) + 1; if (off < subprog_start || off >= subprog_end) { verbose(env, "jump out of range from insn %d to %d\n", i, off); @@ -2889,7 +2891,8 @@ static int check_subprogs(struct bpf_verifier_env *env) */ if (code != (BPF_JMP | BPF_EXIT) && code != (BPF_JMP32 | BPF_JA) && - code != (BPF_JMP | BPF_JA)) { + code != (BPF_JMP | BPF_JA) && + !insn_is_gotox(&insn[i])) { verbose(env, "last insn is not an exit or jmp\n"); return -EINVAL; } @@ -9261,6 +9264,16 @@ static int btf_check_func_arg_match(struct bpf_verifier_env *env, int subprog, return ret; if (check_mem_reg(env, reg, argno, arg->mem_size)) return -EINVAL; + /* + * PTR_TO_PACKET get passed as PTR_TO_MEM, preventing + * us from adjusting bounds tracking info. + */ + if ((reg_is_pkt_pointer_any(reg) || reg_is_dynptr_slice_pkt(reg)) && + sub->changes_pkt_data) { + bpf_log(log, "%s is a packet pointer, but func#%d may change packet data\n", + reg_arg_name(env, argno), subprog); + return -EINVAL; + } if (!(arg->arg_type & PTR_MAYBE_NULL) && (type_may_be_null(reg->type) || bpf_register_is_null(reg))) { bpf_log(log, "%s is expected to be non-NULL\n", diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c index 69ff86e03c63..c43f15c9917c 100644 --- a/kernel/cgroup/cpuset.c +++ b/kernel/cgroup/cpuset.c @@ -1552,10 +1552,11 @@ static int remote_partition_enable(struct cpuset *cs, int new_prs, * above it or remote partition root underneath it is not allowed. */ compute_excpus(cs, tmp->new_cpus); - WARN_ON_ONCE(cpumask_intersects(tmp->new_cpus, subpartitions_cpus)); if (!cpumask_intersects(tmp->new_cpus, cpu_active_mask) || cpumask_subset(top_cpuset.effective_cpus, tmp->new_cpus)) return PERR_INVCPUS; + if (cpumask_intersects(tmp->new_cpus, subpartitions_cpus)) + return PERR_NOCPUS; if (((new_prs == PRS_ISOLATED) && !isolated_cpus_can_update(tmp->new_cpus, NULL)) || prstate_housekeeping_conflict(new_prs, tmp->new_cpus)) diff --git a/kernel/cgroup/pids.c b/kernel/cgroup/pids.c index ecbb839d2acb..78cdc0558d0c 100644 --- a/kernel/cgroup/pids.c +++ b/kernel/cgroup/pids.c @@ -253,6 +253,11 @@ static void pids_event(struct pids_cgroup *pids_forking, } if (!cgroup_subsys_on_dfl(pids_cgrp_subsys) || cgrp_dfl_root.flags & CGRP_ROOT_PIDS_LOCAL_EVENTS) { + /* + * pids.events reports the local counter on legacy hierarchies + * and when pids_localevents is enabled. + */ + cgroup_file_notify(&p->events_file); cgroup_file_notify(&p->events_local_file); return; } diff --git a/kernel/events/core.c b/kernel/events/core.c index 60e60809bedc..d8f9efaca989 100644 --- a/kernel/events/core.c +++ b/kernel/events/core.c @@ -3764,6 +3764,9 @@ static void perf_ctx_sched_task_cb(struct perf_event_context *ctx, list_for_each_entry(pmu_ctx, &ctx->pmu_ctx_list, pmu_ctx_entry) { cpc = this_cpc(pmu_ctx->pmu); + if (cpc->task_epc != pmu_ctx) + continue; + if (cpc->sched_cb_usage && pmu_ctx->pmu->sched_task) pmu_ctx->pmu->sched_task(pmu_ctx, task, sched_in); } @@ -3914,7 +3917,7 @@ static void __perf_pmu_sched_task(struct perf_cpu_pmu_context *cpc, perf_ctx_lock(cpuctx, cpuctx->task_ctx); perf_pmu_disable(pmu); - pmu->sched_task(cpc->task_epc, task, sched_in); + pmu->sched_task(&cpc->epc, task, sched_in); perf_pmu_enable(pmu); perf_ctx_unlock(cpuctx, cpuctx->task_ctx); @@ -3924,15 +3927,17 @@ static void perf_pmu_sched_task(struct task_struct *prev, struct task_struct *next, bool sched_in) { - struct perf_cpu_context *cpuctx = this_cpu_ptr(&perf_cpu_context); struct perf_cpu_pmu_context *cpc, *cpc2; - /* cpuctx->task_ctx will be handled in perf_event_context_sched_in/out */ - if (prev == next || cpuctx->task_ctx) + if (prev == next) return; - list_for_each_entry_safe(cpc, cpc2, this_cpu_ptr(&sched_cb_list), sched_cb_entry) + list_for_each_entry_safe(cpc, cpc2, this_cpu_ptr(&sched_cb_list), sched_cb_entry) { + if (cpc->task_epc) + continue; + __perf_pmu_sched_task(cpc, sched_in ? next : prev, sched_in); + } } static void perf_event_switch(struct task_struct *task, @@ -5454,6 +5459,8 @@ attach_task_ctx_data(struct task_struct *task, struct kmem_cache *ctx_cache, } if (refcount_inc_not_zero(&old->refcount)) { + if (global) + old->global = true; free_perf_ctx_data(cd); /* unused */ return 0; } diff --git a/kernel/exit.c b/kernel/exit.c index 0abce8edb28b..200a936f601a 100644 --- a/kernel/exit.c +++ b/kernel/exit.c @@ -559,18 +559,23 @@ void mm_update_next_owner(struct mm_struct *mm) */ static void exit_mm_sched_cache(struct mm_struct *mm) { + struct sched_cache_group *grp; unsigned long fp, sub; if (!current->total_numa_faults) return; /* * No lock protection due to performance considerations. - * Make sure mm->sc_stat.footprint does not become + * Make sure the group footprint does not become * negative. */ - fp = READ_ONCE(mm->sc_stat.footprint); + grp = READ_ONCE(mm->sched_cache_grp); + if (!grp) + return; + + fp = READ_ONCE(grp->footprint); sub = min(fp, current->total_numa_faults); - WRITE_ONCE(mm->sc_stat.footprint, fp - sub); + WRITE_ONCE(grp->footprint, fp - sub); } #else static inline void exit_mm_sched_cache(struct mm_struct *mm) diff --git a/kernel/kprobes.c b/kernel/kprobes.c index 6337da5cab9e..4edd8ca5c657 100644 --- a/kernel/kprobes.c +++ b/kernel/kprobes.c @@ -42,6 +42,7 @@ #include #include #include +#include #include #include @@ -526,7 +527,8 @@ enum { OPTIMIZER_ST_FLUSHING = 2, }; -static DECLARE_COMPLETION(optimizer_completion); +/* Bumped at the end of each kprobe_optimizer() pass, under 'kprobe_mutex' */ +static unsigned long optimizer_passes; #define OPTIMIZE_DELAY 5 @@ -654,9 +656,9 @@ static void kprobe_optimizer(void) do_free_cleaned_kprobes(); } - /* Step 5: Kick optimizer again if needed. But if there is a flush requested, */ - if (completion_done(&optimizer_completion)) - complete(&optimizer_completion); + /* Step 5: Wake up flushers, and kick optimizer again if needed. */ + optimizer_passes++; + wake_up_var_locked(&optimizer_passes, &kprobe_mutex); if (!list_empty(&optimizing_list) || !list_empty(&unoptimizing_list)) kick_kprobe_optimizer(); /*normal kick*/ @@ -708,7 +710,8 @@ static void wait_for_kprobe_optimizer_locked(void) lockdep_assert_held(&kprobe_mutex); while (!list_empty(&optimizing_list) || !list_empty(&unoptimizing_list)) { - init_completion(&optimizer_completion); + unsigned long passes = optimizer_passes; + /* * Set state to OPTIMIZER_ST_FLUSHING and wake up the thread if it's * idle. If it's already kicked, it will see the state change. @@ -717,9 +720,12 @@ static void wait_for_kprobe_optimizer_locked(void) OPTIMIZER_ST_FLUSHING) != OPTIMIZER_ST_FLUSHING) wake_up(&kprobe_optimizer_wait); - mutex_unlock(&kprobe_mutex); - wait_for_completion(&optimizer_completion); - mutex_lock(&kprobe_mutex); + /* + * kprobe_optimizer() holds 'kprobe_mutex' for a whole pass, which + * this drops while sleeping, so a new count means a full pass ran. + */ + wait_var_event_mutex(&optimizer_passes, + optimizer_passes != passes, &kprobe_mutex); } } diff --git a/kernel/power/hibernate.c b/kernel/power/hibernate.c index d2479c69d71a..c13f68ab7f6e 100644 --- a/kernel/power/hibernate.c +++ b/kernel/power/hibernate.c @@ -408,9 +408,18 @@ int hibernation_snapshot(int platform_mode) if (error) goto Close; + error = dpm_prepare(PMSG_FREEZE); + if (error) + goto Complete; + + /* Preallocate image memory before freezing kernel threads and shutting down devices. */ + error = hibernate_preallocate_memory(); + if (error) + goto Complete; + error = freeze_kernel_threads(); if (error) - goto Close; + goto Cleanup; if (hibernation_test(TEST_FREEZER)) { @@ -422,15 +431,6 @@ int hibernation_snapshot(int platform_mode) goto Thaw; } - error = dpm_prepare(PMSG_FREEZE); - if (error) - goto Complete; - - /* Preallocate image memory before shutting down devices. */ - error = hibernate_preallocate_memory(); - if (error) - goto Complete; - console_suspend_all(); pm_restrict_gfp_mask(); @@ -464,10 +464,12 @@ int hibernation_snapshot(int platform_mode) platform_end(platform_mode); return error; - Complete: - dpm_complete(PMSG_RECOVER); Thaw: thaw_kernel_threads(); + Cleanup: + swsusp_free(); + Complete: + dpm_complete(PMSG_RECOVER); goto Close; } diff --git a/kernel/sched/core.c b/kernel/sched/core.c index 3d6c55598726..c104562638e3 100644 --- a/kernel/sched/core.c +++ b/kernel/sched/core.c @@ -5802,7 +5802,7 @@ void sched_tick(void) curr = rq->curr; donor = rq->donor; - psi_account_irqtime(rq, donor, NULL); + psi_account_irqtime(rq, curr, NULL); update_rq_clock(rq); hw_pressure = arch_scale_hw_pressure(cpu_of(rq)); diff --git a/kernel/sched/ext/ext.c b/kernel/sched/ext/ext.c index 906e4f08967f..2d0ab78b0efa 100644 --- a/kernel/sched/ext/ext.c +++ b/kernel/sched/ext/ext.c @@ -2037,7 +2037,12 @@ static void enqueue_task_scx(struct rq *rq, struct task_struct *p, int core_enq_ int sticky_cpu = p->scx.sticky_cpu; u64 enq_flags = core_enq_flags | rq->scx.extra_enq_flags; - if (enq_flags & ENQUEUE_WAKEUP) + /* + * SCX_RQ_IN_WAKEUP promises a task_woken_scx() call once this enqueue + * returns. Only the core's wakeup path delivers one. The flags stashed + * for a remote activation may carry the wakeup bit without it. + */ + if (core_enq_flags & ENQUEUE_WAKEUP) rq->scx.flags |= SCX_RQ_IN_WAKEUP; /* diff --git a/kernel/sched/fair.c b/kernel/sched/fair.c index bb8a5f358ee1..db34eb74804d 100644 --- a/kernel/sched/fair.c +++ b/kernel/sched/fair.c @@ -1453,12 +1453,17 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) return true; if (static_branch_likely(&sched_numa_balancing)) { + struct sched_cache_group *grp = READ_ONCE(mm->sched_cache_grp); + + if (!grp) + return true; + /* * TBD: RDT exclusive LLC ways reserved should be * excluded. */ llc = sd->llc_bytes; - footprint = READ_ONCE(mm->sc_stat.footprint); + footprint = READ_ONCE(grp->footprint); /* * Scale the LLC size by 256*llc_aggr_tolerance @@ -1490,6 +1495,7 @@ static bool exceed_llc_capacity(struct mm_struct *mm, int cpu) static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p, int cpu) { + struct sched_cache_group *grp; int scale; if (get_nr_threads(p) <= 1) @@ -1503,7 +1509,11 @@ static bool invalid_llc_nr(struct mm_struct *mm, struct task_struct *p, if (scale == INT_MAX) return false; - return !fits_capacity((mm->sc_stat.nr_running_avg * cpu_smt_num_threads), + grp = READ_ONCE(mm->sched_cache_grp); + if (!grp) + return true; + + return !fits_capacity((READ_ONCE(grp->nr_running_avg) * cpu_smt_num_threads), (scale * per_cpu(sd_llc_size, cpu))); } @@ -1580,12 +1590,20 @@ static void account_llc_dequeue(struct rq *rq, struct task_struct *p) } } -void mm_init_sched(struct mm_struct *mm, - struct sched_cache_time __percpu *_pcpu_sched) +int mm_init_sched(struct mm_struct *mm, + struct sched_cache_time __percpu *_pcpu_sched) { + struct sched_cache_group *grp; unsigned long epoch = 0; int i; + grp = kzalloc_obj(*grp); + if (!grp) { + free_percpu(_pcpu_sched); + mm->sched_cache_grp = NULL; + return -ENOMEM; + } + for_each_possible_cpu(i) { struct sched_cache_time *pcpu_sched = per_cpu_ptr(_pcpu_sched, i); struct rq *rq = cpu_rq(i); @@ -1596,18 +1614,51 @@ void mm_init_sched(struct mm_struct *mm, epoch = rq->cpu_epoch; } - raw_spin_lock_init(&mm->sc_stat.lock); - mm->sc_stat.epoch = epoch; - mm->sc_stat.cpu = -1; - mm->sc_stat.next_scan = jiffies; - mm->sc_stat.nr_running_avg = 0; - mm->sc_stat.footprint = 0; + raw_spin_lock_init(&grp->lock); + grp->epoch = epoch; + grp->cpu = -1; + grp->next_scan = jiffies; + grp->nr_running_avg = 0; + grp->footprint = 0; + refcount_set(&grp->refcnt, 1); /* - * The update to mm->sc_stat should not be reordered - * before initialization to mm's other fields, in case + * The update to grp->pcpu_sched should not be reordered + * before initialization to grp's other fields, in case * the readers may get invalid mm_sched_epoch, etc. */ - smp_store_release(&mm->sc_stat.pcpu_sched, _pcpu_sched); + smp_store_release(&grp->pcpu_sched, _pcpu_sched); + /* + * Publish the group last. Not every reader qualifies it by + * grp->pcpu_sched - can_migrate_llc_task() only checks that the + * pointer is non-NULL before reading grp->footprint and + * grp->nr_running_avg - so a reachable group must already be + * fully initialized. + */ + smp_store_release(&mm->sched_cache_grp, grp); + return 0; +} + +static void sched_cache_group_free_rcu(struct rcu_head *rcu) +{ + struct sched_cache_group *grp = + container_of(rcu, struct sched_cache_group, rcu); + + free_percpu(grp->pcpu_sched); + kfree(grp); +} + +static void sched_cache_group_put(struct sched_cache_group *grp) +{ + if (!grp || !refcount_dec_and_test(&grp->refcnt)) + return; + + call_rcu(&grp->rcu, sched_cache_group_free_rcu); +} + +void mm_destroy_sched(struct mm_struct *mm) +{ + sched_cache_group_put(mm->sched_cache_grp); + mm->sched_cache_grp = NULL; } /* because why would C be fully specified */ @@ -1661,11 +1712,16 @@ static unsigned long fraction_mm_sched(struct rq *rq, static int get_pref_llc(struct task_struct *p, struct mm_struct *mm) { int mm_sched_llc = -1, mm_sched_cpu; + struct sched_cache_group *grp; if (!mm) return -1; - mm_sched_cpu = READ_ONCE(mm->sc_stat.cpu); + grp = READ_ONCE(mm->sched_cache_grp); + if (!grp) + return -1; + + mm_sched_cpu = READ_ONCE(grp->cpu); if (mm_sched_cpu != -1) { mm_sched_llc = llc_id(mm_sched_cpu); @@ -1696,6 +1752,7 @@ static inline void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) { struct sched_cache_time *pcpu_sched; + struct sched_cache_group *grp; struct mm_struct *mm = p->mm; int mm_sched_llc = -1; unsigned long epoch; @@ -1709,10 +1766,14 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) * init_task, kthreads and user thread created * by user_mode_thread() don't have mm. */ - if (!mm || !mm->sc_stat.pcpu_sched) + if (!mm) + return; + + grp = READ_ONCE(mm->sched_cache_grp); + if (!grp || !grp->pcpu_sched) return; - pcpu_sched = per_cpu_ptr(mm->sc_stat.pcpu_sched, cpu_of(rq)); + pcpu_sched = per_cpu_ptr(grp->pcpu_sched, cpu_of(rq)); scoped_guard (raw_spinlock, &rq->cpu_epoch_lock) { __update_mm_sched(rq, pcpu_sched); @@ -1725,11 +1786,11 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) * If this process hasn't hit task_cache_work() for a while invalidate * its preferred state. */ - if ((long)(epoch - READ_ONCE(mm->sc_stat.epoch)) > llc_epoch_affinity_timeout || + if ((long)(epoch - READ_ONCE(grp->epoch)) > llc_epoch_affinity_timeout || invalid_llc_nr(mm, p, cpu_of(rq)) || exceed_llc_capacity(mm, cpu_of(rq))) { - if (READ_ONCE(mm->sc_stat.cpu) != -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(grp->cpu) != -1) + WRITE_ONCE(grp->cpu, -1); } mm_sched_llc = get_pref_llc(p, mm); @@ -1746,30 +1807,35 @@ void account_mm_sched(struct rq *rq, struct task_struct *p, s64 delta_exec) static void task_tick_cache(struct rq *rq, struct task_struct *p) { struct callback_head *work = &p->cache_work; + struct sched_cache_group *grp; struct mm_struct *mm = p->mm; unsigned long epoch; if (!sched_cache_enabled()) return; - if (!mm || p->flags & PF_KTHREAD || - !mm->sc_stat.pcpu_sched) + if (!mm || p->flags & PF_KTHREAD) + return; + + grp = READ_ONCE(mm->sched_cache_grp); + if (!grp || !grp->pcpu_sched) return; epoch = rq->cpu_epoch; /* avoid moving backwards */ - if (time_after_eq(mm->sc_stat.epoch, epoch)) + if (time_after_eq(grp->epoch, epoch)) return; - guard(raw_spinlock)(&mm->sc_stat.lock); + guard(raw_spinlock)(&grp->lock); if (work->next == work) { task_work_add(p, work, TWA_RESUME); - WRITE_ONCE(mm->sc_stat.epoch, epoch); + WRITE_ONCE(grp->epoch, epoch); } } -static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p) +static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p, + struct sched_cache_group *grp) { #ifdef CONFIG_NUMA_BALANCING int cpu, curr_cpu, nid, pref_nid; @@ -1777,7 +1843,7 @@ static void get_scan_cpumasks(cpumask_var_t cpus, struct task_struct *p) if (!static_branch_likely(&sched_numa_balancing)) goto out; - cpu = READ_ONCE(p->mm->sc_stat.cpu); + cpu = READ_ONCE(grp->cpu); if (cpu != -1) nid = cpu_to_node(cpu); curr_cpu = task_cpu(p); @@ -1838,6 +1904,7 @@ static void task_cache_work(struct callback_head *work) unsigned long next_scan, now = jiffies; struct task_struct *p = current, *cur; unsigned long curr_m_a_occ = 0; + struct sched_cache_group *grp; struct mm_struct *mm = p->mm; unsigned long m_a_occ = 0; cpumask_var_t cpus; @@ -1849,12 +1916,16 @@ static void task_cache_work(struct callback_head *work) if (p->flags & PF_EXITING) return; - next_scan = READ_ONCE(mm->sc_stat.next_scan); + grp = READ_ONCE(mm->sched_cache_grp); + if (!grp) + return; + + next_scan = READ_ONCE(grp->next_scan); if (time_before(now, next_scan)) return; /* only 1 thread is allowed to scan */ - if (!try_cmpxchg(&mm->sc_stat.next_scan, &next_scan, + if (!try_cmpxchg(&grp->next_scan, &next_scan, now + max_t(unsigned long, READ_ONCE(llc_epoch_period), 1))) return; @@ -1862,8 +1933,8 @@ static void task_cache_work(struct callback_head *work) curr_cpu = task_cpu(p); if (invalid_llc_nr(mm, p, curr_cpu) || exceed_llc_capacity(mm, curr_cpu)) { - if (READ_ONCE(mm->sc_stat.cpu) != -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(grp->cpu) != -1) + WRITE_ONCE(grp->cpu, -1); return; } @@ -1874,7 +1945,7 @@ static void task_cache_work(struct callback_head *work) scoped_guard (cpus_read_lock) { guard(rcu)(); - get_scan_cpumasks(cpus, p); + get_scan_cpumasks(cpus, p, grp); for_each_cpu(cpu, cpus) { /* XXX sched_cluster_active */ @@ -1887,7 +1958,7 @@ static void task_cache_work(struct callback_head *work) for_each_cpu(i, sched_domain_span(sd)) { occ = fraction_mm_sched(cpu_rq(i), - per_cpu_ptr(mm->sc_stat.pcpu_sched, i)); + per_cpu_ptr(grp->pcpu_sched, i)); a_occ += occ; if (occ > m_occ) { m_occ = occ; @@ -1920,7 +1991,7 @@ static void task_cache_work(struct callback_head *work) m_a_cpu = m_cpu; } - if (llc_id(cpu) == llc_id(READ_ONCE(mm->sc_stat.cpu))) + if (llc_id(cpu) == llc_id(READ_ONCE(grp->cpu))) curr_m_a_occ = a_occ; cpumask_andnot(cpus, cpus, sched_domain_span(sd)); @@ -1929,7 +2000,7 @@ static void task_cache_work(struct callback_head *work) if (m_a_occ > (2 * curr_m_a_occ)) { /* - * Avoid switching sc_stat.cpu too fast. + * Avoid switching sched_cache_grp->cpu too fast. * The reason to choose 2X is because: * 1. It is better to keep the preferred LLC stable, * rather than changing it frequently and cause migrations @@ -1938,10 +2009,10 @@ static void task_cache_work(struct callback_head *work) * 3. 2X is chosen based on test results, as it delivers * the optimal performance gain so far. */ - WRITE_ONCE(mm->sc_stat.cpu, m_a_cpu); + WRITE_ONCE(grp->cpu, m_a_cpu); } - update_avg_scale(&mm->sc_stat.nr_running_avg, nr_running); + update_avg_scale(&grp->nr_running_avg, nr_running); free_cpumask_var(cpus); } @@ -3647,6 +3718,7 @@ static int preferred_group_nid(struct task_struct *p, int nid) static void task_numa_placement(struct task_struct *p) __context_unsafe(/* conditional locking */) { + struct sched_cache_group __maybe_unused *grp; int seq, nid, max_nid = NUMA_NO_NODE; unsigned long max_faults = 0; unsigned long fault_types[2] = { 0, 0 }; @@ -3739,19 +3811,23 @@ static void task_numa_placement(struct task_struct *p) * heuristic and occasional lost updates are tolerable. * * If a task exits, its corresponding footprint must - * be subtracted from the mm->sc_stat.footprint, otherwise - * the mm->sc_stat.footprint will not converge: - * the exiting thread's footprint remains unchanged/undecayed - * in mm->sc_stat.footprint. See exit_mm(). + * be subtracted from the mm->sched_cache_grp->footprint, + * otherwise the mm->sched_cache_grp->footprint will not + * converge: the exiting thread's footprint remains + * unchanged/undecayed in mm->sched_cache_grp->footprint. + * See exit_mm(). * * Lost updates and unsynchronized subtraction * in exit_mm() can cause footprint + diff to * go negative. Clamp to zero to prevent the * unsigned footprint from wrapping. */ - new_fp = (long)READ_ONCE(p->mm->sc_stat.footprint) + diff; - WRITE_ONCE(p->mm->sc_stat.footprint, - max(new_fp, 0L)); + grp = READ_ONCE(p->mm->sched_cache_grp); + if (!grp) + continue; + + new_fp = (long)READ_ONCE(grp->footprint) + diff; + WRITE_ONCE(grp->footprint, max(new_fp, 0L)); #endif } @@ -10582,6 +10658,7 @@ static enum llc_mig can_migrate_llc(int src_cpu, int dst_cpu, static enum llc_mig can_migrate_llc_task(int src_cpu, int dst_cpu, struct task_struct *p) { + struct sched_cache_group *grp; struct mm_struct *mm; bool to_pref; int cpu; @@ -10590,15 +10667,19 @@ static enum llc_mig can_migrate_llc_task(int src_cpu, int dst_cpu, if (!mm) return mig_unrestricted; - cpu = READ_ONCE(mm->sc_stat.cpu); + grp = READ_ONCE(mm->sched_cache_grp); + if (!grp) + return mig_unrestricted; + + cpu = READ_ONCE(grp->cpu); if (cpu < 0 || cpus_share_cache(src_cpu, dst_cpu)) return mig_unrestricted; /* skip cache aware load balance for too many threads */ if (invalid_llc_nr(mm, p, dst_cpu) || exceed_llc_capacity(mm, dst_cpu)) { - if (READ_ONCE(mm->sc_stat.cpu) != -1) - WRITE_ONCE(mm->sc_stat.cpu, -1); + if (READ_ONCE(grp->cpu) != -1) + WRITE_ONCE(grp->cpu, -1); return mig_unrestricted; } diff --git a/kernel/sched/topology.c b/kernel/sched/topology.c index 622e2e01974c..70b63224a230 100644 --- a/kernel/sched/topology.c +++ b/kernel/sched/topology.c @@ -985,8 +985,8 @@ void sched_cache_active_set(void) } /* - * Update the bottom sched_domain's llc_bytes for @cpu and all its - * LLC siblings. Called from cacheinfo_cpu_online() or + * Update the bottom sched_domain's llc_bytes for @cpus sharing a physical + * LLC. Called from cacheinfo_cpu_online() or * cacheinfo_cpu_pre_down() with cpu hotplug lock held. * * Note: get_effective_llc_bytes() returns 0 on PowerPC. @@ -996,17 +996,13 @@ void sched_cache_active_set(void) * and does not populates the per-CPU struct cpu_cacheinfo array * that get_cpu_cacheinfo_llc() reads. */ -void sched_update_llc_bytes(unsigned int cpu) +void sched_update_llc_bytes(const struct cpumask *cpus) { struct sched_domain *sd, *sdp; unsigned int i; sched_domains_mutex_lock(); - sdp = rcu_dereference_sched_domain(per_cpu(sd_llc, cpu)); - if (!sdp) - goto unlock; - /* * ci->shared_cpu_map is built incrementally as CPUs come * online, so the first CPU in an LLC initially sees @@ -1014,14 +1010,22 @@ void sched_update_llc_bytes(unsigned int cpu) * get_effective_llc_bytes(). Re-evaluating every LLC * sibling on each online event corrects this once the full * shared_cpu_map is known. + * + * The departing CPU's domains have already been detached when + * cacheinfo removes it. Use the surviving cache siblings instead. + * They may belong to different cpuset partitions, so use each CPU's + * own LLC domain to scale its share of the physical cache. */ - for_each_cpu(i, sched_domain_span(sdp)) { + for_each_cpu(i, cpus) { + sdp = rcu_dereference_sched_domain(per_cpu(sd_llc, i)); + if (!sdp) + continue; + sd = rcu_dereference_sched_domain(cpu_rq(i)->sd); if (sd) sd->llc_bytes = get_effective_llc_bytes(i, sdp); } -unlock: sched_domains_mutex_unlock(); } diff --git a/kernel/trace/fprobe.c b/kernel/trace/fprobe.c index f681015413b8..2a1d65754b04 100644 --- a/kernel/trace/fprobe.c +++ b/kernel/trace/fprobe.c @@ -171,6 +171,11 @@ static inline bool write_fprobe_header(unsigned long *stack, static inline void read_fprobe_header(unsigned long *stack, struct fprobe **fp, unsigned int *size_words) { + if (!*stack) { + *fp = NULL; + *size_words = 0; + return; + } *fp = arch_decode_fprobe_header_fp(*stack); *size_words = arch_decode_fprobe_header_size(*stack); } @@ -203,6 +208,12 @@ static inline void read_fprobe_header(unsigned long *stack, { struct __fprobe_header *fph = (struct __fprobe_header *)stack; + if (!*stack) { + *fp = NULL; + *size_words = 0; + return; + } + *fp = fph->fp; *size_words = fph->size_words; } @@ -642,6 +653,10 @@ static int fprobe_fgraph_entry(struct ftrace_graph_ent *trace, struct fgraph_ops } } + /* Terminate the list, fgraph_reserve_data() does not clear it. */ + if (used && used < reserved_words) + fgraph_data[used] = 0; + /* If any exit_handler is set, data must be used. */ return used != 0; } diff --git a/kernel/workqueue.c b/kernel/workqueue.c index 36aaeb38e217..53608ac01fab 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -3844,7 +3844,7 @@ static void check_flush_dependency(struct workqueue_struct *target_wq, WARN_ONCE(current->flags & PF_MEMALLOC, "workqueue: PF_MEMALLOC task %d(%s) is flushing !WQ_MEM_RECLAIM %s:%ps", current->pid, current->comm, target_wq->name, target_func); - WARN_ONCE(worker && ((worker->current_pwq->wq->flags & + WARN_ONCE(worker && worker->current_pwq && ((worker->current_pwq->wq->flags & (WQ_MEM_RECLAIM | __WQ_LEGACY)) == WQ_MEM_RECLAIM), "workqueue: WQ_MEM_RECLAIM %s:%ps is flushing !WQ_MEM_RECLAIM %s:%ps", worker->current_pwq->wq->name, worker->current_func, diff --git a/mm/backing-dev.c b/mm/backing-dev.c index cecbcf9060a6..18e999053bae 100644 --- a/mm/backing-dev.c +++ b/mm/backing-dev.c @@ -910,8 +910,9 @@ static void cleanup_offline_cgwbs_workfn(struct work_struct *work) continue; spin_unlock_irq(&cgwb_lock); - while (cleanup_offline_cgwb(wb)) - cond_resched(); + do { + cond_resched_tasks_rcu_qs(); + } while (cleanup_offline_cgwb(wb)); spin_lock_irq(&cgwb_lock); wb_put(wb); diff --git a/mm/damon/core.c b/mm/damon/core.c index 372ca1161c57..c66fd59eb8ff 100644 --- a/mm/damon/core.c +++ b/mm/damon/core.c @@ -2145,36 +2145,39 @@ static bool damos_skip_charged_region(struct damon_target *t, { struct damos_quota *quota = &s->quota; unsigned long sz_to_skip; + bool skip = false; /* Skip previously charged regions */ if (quota->charge_target_from) { if (t != quota->charge_target_from) return true; - if (r == damon_last_region(t)) { - quota->charge_target_from = NULL; - quota->charge_addr_from = 0; - return true; - } if (quota->charge_addr_from && - r->ar.end <= quota->charge_addr_from) - return true; + r->ar.end <= quota->charge_addr_from) { + skip = true; + goto out; + } if (quota->charge_addr_from && r->ar.start < quota->charge_addr_from) { sz_to_skip = ALIGN_DOWN(quota->charge_addr_from - r->ar.start, min_region_sz); if (!sz_to_skip) { - if (damon_sz_region(r) <= min_region_sz) - return true; + if (damon_sz_region(r) <= min_region_sz) { + skip = true; + goto out; + } sz_to_skip = min_region_sz; } damon_split_region_at(t, r, sz_to_skip); - return true; + skip = true; } + } +out: + if (r == damon_last_region(t)) { quota->charge_target_from = NULL; quota->charge_addr_from = 0; } - return false; + return skip; } static void damos_update_stat(struct damos *s, @@ -2893,6 +2896,7 @@ static void damos_set_effective_quota(struct damon_ctx *ctx, struct damos *s) struct damos_quota *quota = &s->quota; unsigned long throughput; unsigned long esz = ULONG_MAX; + unsigned long esz_time; if (!quota->ms && list_empty("a->goals)) { quota->esz = quota->sz; @@ -2913,8 +2917,8 @@ static void damos_set_effective_quota(struct damon_ctx *ctx, struct damos *s) 1000000, quota->total_charged_ns); else throughput = PAGE_SIZE * 1024; - esz = min(throughput * quota->ms, esz); - esz = max(ctx->min_region_sz, esz); + esz_time = max(throughput * quota->ms, ctx->min_region_sz); + esz = min(esz_time, esz); } if (quota->sz && quota->sz < esz) @@ -3043,8 +3047,15 @@ static void kdamond_apply_schemes(struct damon_ctx *c) max_region_sz = damon_region_sz_limit(c); mutex_lock(&c->walk_control_lock); damon_for_each_target(t, c) { - if (c->ops.target_valid && c->ops.target_valid(t) == false) + if (c->ops.target_valid && c->ops.target_valid(t) == false) { + damon_for_each_scheme(s, c) { + if (s->quota.charge_target_from != t) + continue; + s->quota.charge_target_from = NULL; + s->quota.charge_addr_from = 0; + } continue; + } damos_apply_target(c, t, max_region_sz); } diff --git a/mm/damon/ops-common.c b/mm/damon/ops-common.c index 6a969b1d2987..58ccb14f80b3 100644 --- a/mm/damon/ops-common.c +++ b/mm/damon/ops-common.c @@ -61,7 +61,12 @@ void damon_ptep_mkold(pte_t *pte, struct vm_area_struct *vma, unsigned long addr * device aspects. */ if (likely(pte_present(pteval))) - young |= ptep_test_and_clear_young(vma, addr, pte); + /* + * Arch implementation of ptep_test_and_clear_young() may + * require aligned @addr + */ + young |= ptep_test_and_clear_young(vma, PAGE_ALIGN_DOWN(addr), + pte); young |= mmu_notifier_clear_young(vma->vm_mm, addr, addr + PAGE_SIZE); if (young) folio_set_young(folio); diff --git a/mm/damon/vaddr.c b/mm/damon/vaddr.c index 404785128ffd..fb4f4911ea70 100644 --- a/mm/damon/vaddr.c +++ b/mm/damon/vaddr.c @@ -292,22 +292,29 @@ static int damon_mkold_pmd_entry(pmd_t *pmd, unsigned long addr, } #ifdef CONFIG_HUGETLB_PAGE +static bool damon_hugetlb_ptep_mkold(pte_t *pte, struct mm_struct *mm, + struct vm_area_struct *vma, unsigned long addr, pte_t *entry) +{ + unsigned long psize = huge_page_size(hstate_vma(vma)); + + if (!pte_young(*entry)) + return false; + *entry = huge_ptep_get_and_clear(mm, addr, pte, psize); + *entry = pte_mkold(*entry); + set_huge_pte_at(mm, addr, pte, *entry, psize); + return true; +} + static void damon_hugetlb_mkold(pte_t *pte, struct mm_struct *mm, struct vm_area_struct *vma, unsigned long addr) { bool referenced = false; pte_t entry = huge_ptep_get(mm, addr, pte); struct folio *folio = pfn_folio(pte_pfn(entry)); - unsigned long psize = huge_page_size(hstate_vma(vma)); folio_get(folio); - if (pte_young(entry)) { - referenced = true; - entry = pte_mkold(entry); - set_huge_pte_at(mm, addr, pte, entry, psize); - } - + referenced = damon_hugetlb_ptep_mkold(pte, mm, vma, addr, &entry); if (mmu_notifier_clear_young(mm, addr, addr + huge_page_size(hstate_vma(vma)))) referenced = true; diff --git a/mm/hugetlb.c b/mm/hugetlb.c index ca080806de17..abb15476eff1 100644 --- a/mm/hugetlb.c +++ b/mm/hugetlb.c @@ -1977,6 +1977,15 @@ int dissolve_free_hugetlb_folio(struct folio *folio) struct hstate *h = folio_hstate(folio); bool adjust_surplus = false; + /* + * remove_hugetlb_folio()/update_and_free_hugetlb_folio() bail + * for gigantic hstates without runtime support, so dissolving one + * here would leave it on the free list and, on vmemmap restore + * failure, the add_hugetlb_folio() rollback corrupts that list. + */ + if (hstate_is_gigantic_no_runtime(h)) + goto out; + if (!available_huge_pages(h)) goto out; @@ -5134,18 +5143,21 @@ int move_hugetlb_page_tables(struct vm_area_struct *vma, hugetlb_vma_lock_write(vma); i_mmap_lock_write(mapping); for (; old_addr < old_end; old_addr += sz, new_addr += sz) { + const unsigned long offset_to_last_entry = + (old_addr | last_addr_mask) - old_addr; + src_pte = hugetlb_walk(vma, old_addr, sz); if (!src_pte) { - old_addr |= last_addr_mask; - new_addr |= last_addr_mask; + old_addr += offset_to_last_entry; + new_addr += offset_to_last_entry; continue; } if (huge_pte_none(huge_ptep_get(mm, old_addr, src_pte))) continue; if (huge_pmd_unshare(&tlb, vma, old_addr, src_pte)) { - old_addr |= last_addr_mask; - new_addr |= last_addr_mask; + old_addr += offset_to_last_entry; + new_addr += offset_to_last_entry; continue; } diff --git a/mm/rmap.c b/mm/rmap.c index 143cf80d95f4..acdadb470279 100644 --- a/mm/rmap.c +++ b/mm/rmap.c @@ -209,7 +209,11 @@ int __anon_vma_prepare(struct vm_area_struct *vma) /* page_table_lock to protect against threads */ spin_lock(&mm->page_table_lock); if (likely(!vma->anon_vma)) { - vma->anon_vma = anon_vma; + /* + * Make anon_vma fields visible before anon_vma is published. + * Paired with an address dependency in reusable_anon_vma(). + */ + smp_store_release(&vma->anon_vma, anon_vma); anon_vma_chain_assign(vma, avc, anon_vma); anon_vma_interval_tree_insert(avc, &anon_vma->rb_root); anon_vma->num_active_vmas++; diff --git a/mm/vma.c b/mm/vma.c index 98c163431e40..638ec4793798 100644 --- a/mm/vma.c +++ b/mm/vma.c @@ -1995,6 +1995,13 @@ static int anon_vma_compatible(struct vm_area_struct *a, struct vm_area_struct * * acceptable for merging, so we can do all of this optimistically. But * we do that READ_ONCE() to make sure that we never re-load the pointer. * + * The READ_ONCE() establishes an address dependency between anon_vma and + * any access to its fields, which pairs with the assignment to + * vma->anon_vma performed with release semantics in __anon_vma_prepare(). + * + * This is especially important as anon_vma's are SLAB_TYPESAFE_BY_RCU so + * accessing an uninitialised anon_vma's fields may result in a UAF. + * * IOW: that the "list_is_singular()" test on the anon_vma_chain only * matters for the 'stable anon_vma' case (ie the thing we want to avoid * is to return an anon_vma that is "complex" due to having gone through @@ -2009,6 +2016,7 @@ static struct anon_vma *reusable_anon_vma(struct vm_area_struct *old, struct vm_area_struct *b) { if (anon_vma_compatible(a, b)) { + /* Paired with a memory barrier in __anon_vma_prepare(). */ struct anon_vma *anon_vma = READ_ONCE(old->anon_vma); if (anon_vma && list_is_singular(&old->anon_vma_chain)) diff --git a/net/8021q/vlan_dev.c b/net/8021q/vlan_dev.c index 2859cbac3f26..c949c6a82945 100644 --- a/net/8021q/vlan_dev.c +++ b/net/8021q/vlan_dev.c @@ -55,6 +55,11 @@ static int vlan_dev_hard_header(struct sk_buff *skb, struct net_device *dev, int rc; if (!(vlan->flags & VLAN_FLAG_REORDER_HDR)) { + unsigned int hlen = READ_ONCE(dev->hard_header_len) + + READ_ONCE(dev->needed_headroom); + + if (skb_cow_head(skb, hlen) < 0) + return -ENOMEM; vhdr = skb_push(skb, VLAN_HLEN); vlan_tci = vlan->vlan_id; diff --git a/net/bluetooth/bnep/core.c b/net/bluetooth/bnep/core.c index f7d88c33e23e..ad24d2486665 100644 --- a/net/bluetooth/bnep/core.c +++ b/net/bluetooth/bnep/core.c @@ -270,9 +270,14 @@ static int bnep_rx_extension(struct bnep_session *s, struct sk_buff *skb) BT_DBG("type 0x%x len %u", h->type, h->len); + if (skb->len < h->len) { + err = -EILSEQ; + break; + } + switch (h->type & BNEP_TYPE_MASK) { case BNEP_EXT_CONTROL: - bnep_rx_control(s, skb->data, skb->len); + bnep_rx_control(s, skb->data, h->len); break; default: @@ -373,6 +378,11 @@ static int bnep_rx_frame(struct bnep_session *s, struct sk_buff *skb) goto badframe; } + if ((type & BNEP_TYPE_MASK) == BNEP_CONTROL) { + kfree_skb(skb); + return 0; + } + /* Strip 802.1p header */ if (ntohs(s->eh.h_proto) == ETH_P_8021Q) { if (!skb_pull(skb, 4)) @@ -451,6 +461,11 @@ static int bnep_tx_frame(struct bnep_session *s, struct sk_buff *skb) goto send; } + if (skb->len < ETH_HLEN) { + kfree_skb(skb); + return 0; + } + iv[il++] = (struct kvec) { &type, 1 }; len++; diff --git a/net/bluetooth/bnep/netdev.c b/net/bluetooth/bnep/netdev.c index ee1e39a3daff..b451ef457741 100644 --- a/net/bluetooth/bnep/netdev.c +++ b/net/bluetooth/bnep/netdev.c @@ -166,6 +166,12 @@ static netdev_tx_t bnep_net_xmit(struct sk_buff *skb, BT_DBG("skb %p, dev %p", skb, dev); + if (!pskb_may_pull(skb, ETH_HLEN)) { + dev->stats.tx_dropped++; + kfree_skb(skb); + return NETDEV_TX_OK; + } + #ifdef CONFIG_BT_BNEP_MC_FILTER if (bnep_net_mc_filter(skb, s)) { kfree_skb(skb); @@ -218,7 +224,7 @@ void bnep_net_setup(struct net_device *dev) dev->addr_len = ETH_ALEN; ether_setup(dev); - dev->min_mtu = 0; + dev->min_mtu = ETH_MIN_MTU; dev->max_mtu = ETH_MAX_MTU; dev->priv_flags &= ~IFF_TX_SKB_SHARING; dev->netdev_ops = &bnep_netdev_ops; diff --git a/net/bluetooth/hci_conn.c b/net/bluetooth/hci_conn.c index fa72cf8aaa7a..96195d2fd10f 100644 --- a/net/bluetooth/hci_conn.c +++ b/net/bluetooth/hci_conn.c @@ -2079,6 +2079,8 @@ struct hci_conn *hci_bind_cis(struct hci_dev *hdev, bdaddr_t *dst, cis->conn_timeout = timeout; } + hci_conn_hold(cis); + if (cis->state == BT_CONNECTED) return cis; @@ -2120,7 +2122,6 @@ struct hci_conn *hci_bind_cis(struct hci_dev *hdev, bdaddr_t *dst, return ERR_PTR(-EINVAL); } - hci_conn_hold(cis); cis->state = BT_BOUND; return cis; @@ -2374,10 +2375,13 @@ struct hci_conn *hci_bind_bis(struct hci_dev *hdev, bdaddr_t *dst, __u8 sid, parent = hci_conn_hash_lookup_big(hdev, conn->iso_qos.bcast.big); if (parent && parent != conn) { + hci_conn_hold(parent); link = hci_conn_link(parent, conn); hci_conn_drop(conn); - if (!link) + if (!link) { + hci_conn_drop(parent); return ERR_PTR(-ENOLINK); + } } return conn; @@ -2497,6 +2501,12 @@ struct hci_conn *hci_connect_cis(struct hci_dev *hdev, bdaddr_t *dst, return cis; } + /* The existing link already owns the hold on its parent. */ + if (cis->link) { + hci_conn_drop(le); + return cis; + } + link = hci_conn_link(le, cis); hci_conn_drop(cis); if (!link) { diff --git a/net/bluetooth/hci_sock.c b/net/bluetooth/hci_sock.c index 070ca388f9ac..6d56c77741e1 100644 --- a/net/bluetooth/hci_sock.c +++ b/net/bluetooth/hci_sock.c @@ -164,6 +164,7 @@ static bool is_filtered_packet(struct sock *sk, struct sk_buff *skb) { struct hci_filter *flt; int flt_type, flt_event; + u8 event; /* Apply filter */ flt = &hci_pi(sk)->filter; @@ -177,7 +178,11 @@ static bool is_filtered_packet(struct sock *sk, struct sk_buff *skb) if (hci_skb_pkt_type(skb) != HCI_EVENT_PKT) return false; - flt_event = (*(__u8 *)skb->data & HCI_FLT_EVENT_BITS); + if (skb->len < 1) + return true; + + event = *(__u8 *)skb->data; + flt_event = event & HCI_FLT_EVENT_BITS; if (!hci_test_bit(flt_event, &flt->event_mask)) return true; @@ -186,11 +191,17 @@ static bool is_filtered_packet(struct sock *sk, struct sk_buff *skb) if (!flt->opcode) return false; - if (flt_event == HCI_EV_CMD_COMPLETE && + if (event == HCI_EV_CMD_COMPLETE && skb->len < 5) + return true; + + if (event == HCI_EV_CMD_COMPLETE && flt->opcode != get_unaligned((__le16 *)(skb->data + 3))) return true; - if (flt_event == HCI_EV_CMD_STATUS && + if (event == HCI_EV_CMD_STATUS && skb->len < 6) + return true; + + if (event == HCI_EV_CMD_STATUS && flt->opcode != get_unaligned((__le16 *)(skb->data + 4))) return true; @@ -1881,7 +1892,8 @@ static int hci_sock_sendmsg(struct socket *sock, struct msghdr *msg, u16 ocf = hci_opcode_ocf(opcode); if (((ogf > HCI_SFLT_MAX_OGF) || - !hci_test_bit(ocf & HCI_FLT_OCF_BITS, + (ocf > HCI_FLT_OCF_BITS) || + !hci_test_bit(ocf, &hci_sec_filter.ocf_mask[ogf])) && !capable(CAP_NET_RAW)) { err = -EPERM; diff --git a/net/bluetooth/iso.c b/net/bluetooth/iso.c index eb99653f33f9..7657c2a0abbf 100644 --- a/net/bluetooth/iso.c +++ b/net/bluetooth/iso.c @@ -496,6 +496,7 @@ static int iso_connect_cis(struct sock *sk) struct hci_dev *hdev; bdaddr_t src, dst; u8 src_type; + bool already_attached; int err; lock_sock(sk); @@ -568,8 +569,14 @@ static int iso_connect_cis(struct sock *sk) goto unlock; } + iso_conn_lock(conn); + already_attached = iso_pi(sk)->conn == conn && conn->sk == sk; + iso_conn_unlock(conn); + err = iso_chan_add(conn, sk, NULL); iso_conn_put(conn); + if (already_attached || err == -EBUSY) + hci_conn_drop(hcon); if (err) goto unlock; diff --git a/net/bluetooth/l2cap_core.c b/net/bluetooth/l2cap_core.c index 800f7517bfeb..9a1b0ae07679 100644 --- a/net/bluetooth/l2cap_core.c +++ b/net/bluetooth/l2cap_core.c @@ -6701,9 +6701,17 @@ static int l2cap_stream_rx(struct l2cap_chan *chan, struct l2cap_ctrl *control, static int l2cap_data_rcv(struct l2cap_chan *chan, struct sk_buff *skb) { struct l2cap_ctrl *control = &bt_cb(skb)->l2cap; - u16 len; + u16 len, min_len; u8 event; + min_len = test_bit(FLAG_EXT_CTRL, &chan->flags) ? + L2CAP_EXT_CTRL_SIZE : L2CAP_ENH_CTRL_SIZE; + if (chan->fcs == L2CAP_FCS_CRC16) + min_len += L2CAP_FCS_SIZE; + + if (skb->len < min_len) + goto drop; + __unpack_control(chan, skb); len = skb->len; diff --git a/net/bluetooth/mgmt.c b/net/bluetooth/mgmt.c index deced8a5bc2d..2a177cd17a9e 100644 --- a/net/bluetooth/mgmt.c +++ b/net/bluetooth/mgmt.c @@ -496,13 +496,9 @@ static int read_unconf_index_list(struct sock *sk, struct hci_dev *hdev, read_lock(&hci_dev_list_lock); - count = 0; - list_for_each_entry(d, &hci_dev_list, list) { - if (hci_dev_test_flag(d, HCI_UNCONFIGURED)) - count++; - } + count = list_count_nodes(&hci_dev_list); - rp_len = sizeof(*rp) + (2 * count); + rp_len = sizeof(*rp) + (sizeof(__le16) * count); rp = kmalloc(rp_len, GFP_ATOMIC); if (!rp) { read_unlock(&hci_dev_list_lock); @@ -2310,6 +2306,8 @@ static void mesh_send_start_complete(struct hci_dev *hdev, void *data, int err) hci_dev_clear_flag(hdev, HCI_MESH_SENDING); /* Send Complete Error Code for handle */ mesh_send_complete(hdev, mesh_tx, false); + if (err != -ECANCELED) + mesh_next(hdev, NULL, 0); return; } @@ -2419,19 +2417,28 @@ static int send_cancel(struct hci_dev *hdev, void *data) do { mesh_tx = mgmt_mesh_next(hdev, cmd->sk); - if (mesh_tx) - mesh_send_complete(hdev, mesh_tx, false); + if (mesh_tx) { + if (!hci_cmd_sync_dequeue(hdev, mesh_send_sync, + mesh_tx, NULL)) + mesh_send_complete(hdev, mesh_tx, false); + } } while (mesh_tx); } else { mesh_tx = mgmt_mesh_find(hdev, cancel->handle); - if (mesh_tx && mesh_tx->sk == cmd->sk) - mesh_send_complete(hdev, mesh_tx, false); + if (mesh_tx && mesh_tx->sk == cmd->sk) { + if (!hci_cmd_sync_dequeue(hdev, mesh_send_sync, + mesh_tx, NULL)) + mesh_send_complete(hdev, mesh_tx, false); + } } mgmt_cmd_complete(cmd->sk, hdev->id, MGMT_OP_MESH_SEND_CANCEL, 0, NULL, 0); + if (!hci_dev_test_flag(hdev, HCI_MESH_SENDING)) + mesh_next(hdev, NULL, 0); + return 0; } diff --git a/net/bluetooth/rfcomm/core.c b/net/bluetooth/rfcomm/core.c index 63fa0f542ccf..3f757198051a 100644 --- a/net/bluetooth/rfcomm/core.c +++ b/net/bluetooth/rfcomm/core.c @@ -1817,7 +1817,8 @@ static struct rfcomm_session *rfcomm_recv_frame(struct rfcomm_session *s, return s; } - if (skb->len < sizeof(*hdr) + 1) { + if (skb->len < sizeof(*hdr) + 1 || + (!__test_ea(hdr->len) && skb->len < sizeof(*hdr) + 2)) { kfree_skb(skb); return s; } diff --git a/net/bluetooth/rfcomm/sock.c b/net/bluetooth/rfcomm/sock.c index 24b42483b3c0..d2855f9b3abc 100644 --- a/net/bluetooth/rfcomm/sock.c +++ b/net/bluetooth/rfcomm/sock.c @@ -785,8 +785,10 @@ static int rfcomm_sock_getsockopt_old(struct socket *sock, int optname, break; case RFCOMM_CONNINFO: - if (sk->sk_state != BT_CONNECTED && - !rfcomm_pi(sk)->dlc->defer_setup) { + if ((sk->sk_state != BT_CONNECTED && + !(sk->sk_state == BT_CONNECT2 && + rfcomm_pi(sk)->dlc->defer_setup)) || + !rfcomm_pi(sk)->dlc->session) { err = -ENOTCONN; break; } diff --git a/net/bluetooth/smp.c b/net/bluetooth/smp.c index c4470958b0d5..14bb2c82a562 100644 --- a/net/bluetooth/smp.c +++ b/net/bluetooth/smp.c @@ -2268,6 +2268,23 @@ static u8 smp_cmd_security_req(struct l2cap_conn *conn, struct sk_buff *skb) bt_dev_dbg(hdev, "conn %p", conn); + /* SMP over BR/EDR only covers cross-transport key derivation; the + * Security Request procedure has no BR/EDR counterpart. Reject it + * here, otherwise smp_ltk_encrypt() finds the peer's LE LTK + * (ADDR_LE_DEV_PUBLIC and BDADDR_BREDR are both 0) and issues + * HCI_OP_LE_START_ENC on the ACL handle, which the controller + * rejects and hci_cs_le_start_enc() turns into a disconnect. Reply + * without smp_failure(): this is not an authentication failure, and + * MGMT_EV_AUTH_FAILED would make bluetoothd drop the device. + */ + if (hcon->type != LE_LINK) { + u8 reason = SMP_CMD_NOTSUPP; + + smp_send_cmd(conn, SMP_CMD_PAIRING_FAIL, sizeof(reason), + &reason); + return 0; + } + if (skb->len < sizeof(*rp)) return SMP_INVALID_PARAMS; diff --git a/net/bridge/br_mdb.c b/net/bridge/br_mdb.c index e0c7020b12f5..a01bd280c722 100644 --- a/net/bridge/br_mdb.c +++ b/net/bridge/br_mdb.c @@ -1523,6 +1523,8 @@ static void br_mdb_flush_pgs(struct net_bridge *br, } br_multicast_del_pg(mp, p, pp); + /* br_multicast_del_pg() can remove other groups from this list. */ + pp = &mp->ports; } } diff --git a/net/bridge/br_stp_bpdu.c b/net/bridge/br_stp_bpdu.c index 74ec42ba1e7d..21d092f5acbb 100644 --- a/net/bridge/br_stp_bpdu.c +++ b/net/bridge/br_stp_bpdu.c @@ -52,7 +52,10 @@ static void br_send_bpdu(struct net_bridge_port *p, LLC_SAP_BSPAN, LLC_PDU_CMD); llc_pdu_init_as_ui_cmd(skb); - llc_mac_hdr_init(skb, p->dev->dev_addr, p->br->group_addr); + if (llc_mac_hdr_init(skb, p->dev->dev_addr, p->br->group_addr)) { + kfree_skb(skb); + return; + } skb_reset_mac_header(skb); diff --git a/net/core/dev.c b/net/core/dev.c index 65cdaf0c81f7..55e1f4045d2c 100644 --- a/net/core/dev.c +++ b/net/core/dev.c @@ -2897,7 +2897,7 @@ int __netif_set_xps_queue(struct net_device *dev, const unsigned long *mask, dev = netdev_get_tx_queue(dev, index)->sb_dev ? : dev; tc = netdev_txq_to_tc(dev, index); - if (tc < 0) + if (tc < 0 || tc >= num_tc) return -EINVAL; } @@ -5370,7 +5370,8 @@ void kick_defer_list_purge(unsigned int cpu) backlog_unlock_irq_restore(sd, flags); } else if (!cmpxchg(&sd->defer_ipi_scheduled, 0, 1)) { - smp_call_function_single_async(cpu, &sd->defer_csd); + if (smp_call_function_single_async(cpu, &sd->defer_csd)) + WRITE_ONCE(sd->defer_ipi_scheduled, 0); } } @@ -6894,25 +6895,35 @@ bool napi_complete_done(struct napi_struct *n, int work_done) } EXPORT_SYMBOL(napi_complete_done); -static void skb_defer_free_flush(void) +static void __skb_defer_free_flush(struct skb_defer_node *sdn, int budget) { struct llist_node *free_list; struct sk_buff *skb, *next; + + if (llist_empty(&sdn->defer_list)) + return; + atomic_long_set(&sdn->defer_count, 0); + free_list = llist_del_all(&sdn->defer_list); + + llist_for_each_entry_safe(skb, next, free_list, ll_node) { + prefetch(next); + napi_consume_skb(skb, budget); + } +} + +void skb_defer_node_flush(struct skb_defer_node *sdn) +{ + __skb_defer_free_flush(sdn, 0); +} + +static void skb_defer_free_flush(void) +{ struct skb_defer_node *sdn; int node; for_each_node(node) { sdn = this_cpu_ptr(net_hotdata.skb_defer_nodes) + node; - - if (llist_empty(&sdn->defer_list)) - continue; - atomic_long_set(&sdn->defer_count, 0); - free_list = llist_del_all(&sdn->defer_list); - - llist_for_each_entry_safe(skb, next, free_list, ll_node) { - prefetch(next); - napi_consume_skb(skb, 1); - } + __skb_defer_free_flush(sdn, 1); } } @@ -12773,6 +12784,7 @@ static int dev_cpu_dead(unsigned int oldcpu) struct sk_buff **list_skb; struct sk_buff *skb; unsigned int cpu; + int node; struct softnet_data *sd, *oldsd, *remsd = NULL; local_irq_disable(); @@ -12833,6 +12845,17 @@ static int dev_cpu_dead(unsigned int oldcpu) rps_input_queue_head_incr(oldsd); } + for_each_node(node) + skb_defer_node_flush(per_cpu_ptr(net_hotdata.skb_defer_nodes, + oldcpu) + node); + node = cpu_to_node(oldcpu); + if (node_possible(node) && + !cpumask_intersects(cpumask_of_node(node), cpu_online_mask)) { + for_each_possible_cpu(cpu) + skb_defer_node_flush(per_cpu_ptr(net_hotdata.skb_defer_nodes, + cpu) + node); + } + return 0; } diff --git a/net/core/dev.h b/net/core/dev.h index b757faead4d1..04fb0e9a571e 100644 --- a/net/core/dev.h +++ b/net/core/dev.h @@ -399,6 +399,8 @@ static inline void napi_assert_will_not_race(const struct napi_struct *napi) WARN_ON(READ_ONCE(napi->list_owner) != -1); } +struct skb_defer_node; +void skb_defer_node_flush(struct skb_defer_node *sdn); void kick_defer_list_purge(unsigned int cpu); int dev_set_hwtstamp_phylib(struct net_device *dev, diff --git a/net/core/dev_ioctl.c b/net/core/dev_ioctl.c index a320e264eaaf..164643140a52 100644 --- a/net/core/dev_ioctl.c +++ b/net/core/dev_ioctl.c @@ -276,19 +276,18 @@ int dev_get_hwtstamp_phylib(struct net_device *dev, if (phy_is_default_hwtstamp(dev->phydev)) return phy_hwtstamp_get(dev->phydev, cfg); + if (!dev->netdev_ops->ndo_hwtstamp_get) + return -EOPNOTSUPP; + return dev->netdev_ops->ndo_hwtstamp_get(dev, cfg); } static int dev_get_hwtstamp(struct net_device *dev, struct ifreq *ifr) { - const struct net_device_ops *ops = dev->netdev_ops; struct kernel_hwtstamp_config kernel_cfg = {}; struct hwtstamp_config cfg; int err; - if (!ops->ndo_hwtstamp_get) - return -EOPNOTSUPP; - if (!netif_device_present(dev)) return -ENODEV; @@ -359,12 +358,18 @@ int dev_set_hwtstamp_phylib(struct net_device *dev, cfg->source = phy_ts ? HWTSTAMP_SOURCE_PHYLIB : HWTSTAMP_SOURCE_NETDEV; if (phy_ts && dev->see_all_hwtstamp_requests) { + if (!ops->ndo_hwtstamp_get) + return -EOPNOTSUPP; + err = ops->ndo_hwtstamp_get(dev, &old_cfg); if (err) return err; } if (!phy_ts || dev->see_all_hwtstamp_requests) { + if (!ops->ndo_hwtstamp_set) + return -EOPNOTSUPP; + err = ops->ndo_hwtstamp_set(dev, cfg, extack); if (err) { if (extack->_msg) @@ -390,7 +395,6 @@ int dev_set_hwtstamp_phylib(struct net_device *dev, static int dev_set_hwtstamp(struct net_device *dev, struct ifreq *ifr) { - const struct net_device_ops *ops = dev->netdev_ops; struct kernel_hwtstamp_config kernel_cfg = {}; struct netlink_ext_ack extack = {}; struct hwtstamp_config cfg; @@ -413,9 +417,6 @@ static int dev_set_hwtstamp(struct net_device *dev, struct ifreq *ifr) return err; } - if (!ops->ndo_hwtstamp_set) - return -EOPNOTSUPP; - if (!netif_device_present(dev)) return -ENODEV; @@ -441,15 +442,11 @@ static int dev_set_hwtstamp(struct net_device *dev, struct ifreq *ifr) int generic_hwtstamp_get_lower(struct net_device *dev, struct kernel_hwtstamp_config *kernel_cfg) { - const struct net_device_ops *ops = dev->netdev_ops; int err; if (!netif_device_present(dev)) return -ENODEV; - if (!ops->ndo_hwtstamp_get) - return -EOPNOTSUPP; - netdev_lock_ops(dev); err = dev_get_hwtstamp_phylib(dev, kernel_cfg); netdev_unlock_ops(dev); @@ -462,15 +459,11 @@ int generic_hwtstamp_set_lower(struct net_device *dev, struct kernel_hwtstamp_config *kernel_cfg, struct netlink_ext_ack *extack) { - const struct net_device_ops *ops = dev->netdev_ops; int err; if (!netif_device_present(dev)) return -ENODEV; - if (!ops->ndo_hwtstamp_set) - return -EOPNOTSUPP; - netdev_lock_ops(dev); err = dev_set_hwtstamp_phylib(dev, kernel_cfg, extack); netdev_unlock_ops(dev); diff --git a/net/core/filter.c b/net/core/filter.c index 1e80a52ef86d..1acc53dd04ef 100644 --- a/net/core/filter.c +++ b/net/core/filter.c @@ -3871,12 +3871,6 @@ static u32 __bpf_skb_min_len(const struct sk_buff *skb) if (offset > 0) min_len = offset; } - if (skb->ip_summed == CHECKSUM_PARTIAL) { - offset = skb_checksum_start_offset(skb) + - skb->csum_offset + sizeof(__sum16); - if (offset > 0) - min_len = offset; - } return min_len; } @@ -3893,6 +3887,11 @@ static int bpf_skb_grow_rcsum(struct sk_buff *skb, unsigned int new_len) static int bpf_skb_trim_rcsum(struct sk_buff *skb, unsigned int new_len) { + if (skb->ip_summed == CHECKSUM_PARTIAL && + new_len < skb_checksum_start_offset(skb) + skb->csum_offset + + sizeof(__sum16)) + skb->ip_summed = CHECKSUM_NONE; + return __skb_trim_rcsum(skb, new_len); } @@ -8882,6 +8881,8 @@ static const struct bpf_func_proto * lwt_seg6local_func_proto(enum bpf_func_id func_id, const struct bpf_prog *prog) { switch (func_id) { + case BPF_FUNC_skb_pull_data: + return NULL; #if IS_ENABLED(CONFIG_IPV6_SEG6_BPF) case BPF_FUNC_lwt_seg6_store_bytes: return &bpf_lwt_seg6_store_bytes_proto; @@ -10402,11 +10403,12 @@ u32 bpf_sock_convert_ctx_access(enum bpf_access_type type, target_size)); *insn++ = BPF_JMP_IMM(BPF_JNE, si->dst_reg, NO_QUEUE_MAPPING, 1); - *insn++ = BPF_MOV64_IMM(si->dst_reg, -1); + *insn++ = BPF_MOV32_IMM(si->dst_reg, -1); #else - *insn++ = BPF_MOV64_IMM(si->dst_reg, -1); - *target_size = 2; + *insn++ = BPF_MOV32_IMM(si->dst_reg, -1); #endif + *target_size = sizeof_field(struct bpf_sock, rx_queue_mapping); + break; } @@ -10942,18 +10944,7 @@ static u32 sock_ops_convert_ctx_access(enum bpf_access_type type, break; case offsetof(struct bpf_sock_ops, rtt_min): - BUILD_BUG_ON(sizeof_field(struct tcp_sock, rtt_min) != - sizeof(struct minmax)); - BUILD_BUG_ON(sizeof(struct minmax) < - sizeof(struct minmax_sample)); - - *insn++ = BPF_LDX_MEM(BPF_FIELD_SIZEOF( - struct bpf_sock_ops_kern, sk), - si->dst_reg, si->src_reg, - offsetof(struct bpf_sock_ops_kern, sk)); - *insn++ = BPF_LDX_MEM(BPF_W, si->dst_reg, si->dst_reg, - offsetof(struct tcp_sock, rtt_min) + - sizeof_field(struct minmax_sample, t)); + SOCK_OPS_GET_FIELD(rtt_min, rtt_min.s[0].v, struct tcp_sock); break; case offsetof(struct bpf_sock_ops, bpf_sock_ops_cb_flags): @@ -12658,8 +12649,9 @@ __bpf_kfunc_start_defs(); * @sock: Pointer to socket to be destroyed * * Return: - * On error, may return EPROTONOSUPPORT, EINVAL. - * EPROTONOSUPPORT if protocol specific destroy handler is not supported. + * On error, may return EOPNOTSUPP, or whatever the protocol specific + * destroy handler returns. + * EOPNOTSUPP if protocol specific destroy handler is not supported. * 0 otherwise */ __bpf_kfunc int bpf_sock_destroy(struct sock_common *sock) @@ -12671,8 +12663,12 @@ __bpf_kfunc int bpf_sock_destroy(struct sock_common *sock) * Supporting protocols will need to acquire sock lock in the BPF context * prior to invoking this kfunc. */ - if (!sk->sk_prot->diag_destroy || (sk->sk_protocol != IPPROTO_TCP && - sk->sk_protocol != IPPROTO_UDP)) + if (!sk->sk_prot->diag_destroy) + return -EOPNOTSUPP; + + if (sk_fullsock(sk) && + sk->sk_protocol != IPPROTO_TCP && + sk->sk_protocol != IPPROTO_UDP) return -EOPNOTSUPP; return sk->sk_prot->diag_destroy(sk, ECONNABORTED); diff --git a/net/core/skbuff.c b/net/core/skbuff.c index c485be081aea..39102b366003 100644 --- a/net/core/skbuff.c +++ b/net/core/skbuff.c @@ -5965,7 +5965,8 @@ static int skb_checksum_setup_ipv6(struct sk_buff *skb, bool recalculate) err = skb_maybe_pull_tail(skb, off + sizeof(struct ipv6_opt_hdr), - MAX_IPV6_HDR_LEN); + off + + sizeof(struct ipv6_opt_hdr)); if (err < 0) goto out; @@ -5980,7 +5981,8 @@ static int skb_checksum_setup_ipv6(struct sk_buff *skb, bool recalculate) err = skb_maybe_pull_tail(skb, off + sizeof(struct ip_auth_hdr), - MAX_IPV6_HDR_LEN); + off + + sizeof(struct ip_auth_hdr)); if (err < 0) goto out; @@ -5995,7 +5997,8 @@ static int skb_checksum_setup_ipv6(struct sk_buff *skb, bool recalculate) err = skb_maybe_pull_tail(skb, off + sizeof(struct frag_hdr), - MAX_IPV6_HDR_LEN); + off + + sizeof(struct frag_hdr)); if (err < 0) goto out; @@ -7345,8 +7348,8 @@ void skb_attempt_defer_free(struct sk_buff *skb) struct skb_defer_node *sdn; unsigned long defer_count; unsigned int defer_max; + int cpu, my_cpu; bool kick; - int cpu; if (static_branch_unlikely(&skb_defer_disable_key)) goto nodefer; @@ -7356,7 +7359,8 @@ void skb_attempt_defer_free(struct sk_buff *skb) goto nodefer; cpu = skb->alloc_cpu; - if (cpu == raw_smp_processor_id() || + my_cpu = raw_smp_processor_id(); + if (cpu == my_cpu || WARN_ON_ONCE(cpu >= nr_cpu_ids) || !cpu_online(cpu)) { nodefer: kfree_skb_napi_cache(skb); @@ -7367,7 +7371,7 @@ nodefer: kfree_skb_napi_cache(skb); DEBUG_NET_WARN_ON_ONCE(skb->destructor); DEBUG_NET_WARN_ON_ONCE(skb_nfct(skb)); - sdn = per_cpu_ptr(net_hotdata.skb_defer_nodes, cpu) + numa_node_id(); + sdn = per_cpu_ptr(net_hotdata.skb_defer_nodes, cpu) + cpu_to_node(my_cpu); defer_max = READ_ONCE(net_hotdata.sysctl_skb_defer_max); defer_count = atomic_long_inc_return(&sdn->defer_count); @@ -7377,6 +7381,11 @@ nodefer: kfree_skb_napi_cache(skb); llist_add(&skb->ll_node, &sdn->defer_list); + if (unlikely(!cpu_online(cpu) || my_cpu != raw_smp_processor_id())) { + skb_defer_node_flush(sdn); + return; + } + /* Send an IPI every time queue reaches half capacity. */ kick = (defer_count - 1) == (defer_max >> 1); diff --git a/net/core/skmsg.c b/net/core/skmsg.c index 2521b643fa05..df385a5a961e 100644 --- a/net/core/skmsg.c +++ b/net/core/skmsg.c @@ -1000,6 +1000,10 @@ static int sk_psock_verdict_apply(struct sk_psock *psock, struct sk_buff *skb, int err = 0; u32 len, off; + if (verdict == __SK_REDIRECT && skb_bpf_ingress(skb) && + skb_bpf_redirect_fetch(skb) == psock->sk) + verdict = __SK_PASS; + switch (verdict) { case __SK_PASS: err = -EIO; diff --git a/net/core/sock_map.c b/net/core/sock_map.c index 9efbd8ca7db8..09d20318a0ea 100644 --- a/net/core/sock_map.c +++ b/net/core/sock_map.c @@ -41,6 +41,7 @@ static struct bpf_map *sock_map_alloc(union bpf_attr *attr) struct bpf_stab *stab; if (attr->max_entries == 0 || + attr->max_entries > INT_MAX || attr->key_size != 4 || (attr->value_size != sizeof(u32) && attr->value_size != sizeof(u64)) || diff --git a/net/ethtool/common.h b/net/ethtool/common.h index 4e5356e26f40..ae32e7fdb563 100644 --- a/net/ethtool/common.h +++ b/net/ethtool/common.h @@ -163,6 +163,8 @@ ethtool_ioctl_needs_rtnl(const struct net_device *dev, u32 ethcmd) return ops->op_needs_rtnl & ETHTOOL_OP_NEEDS_RTNL_RSS; case ETHTOOL_GLINK: return ops->op_needs_rtnl & ETHTOOL_OP_NEEDS_RTNL_GLINK; + case ETHTOOL_TEST: + return ops->op_needs_rtnl & ETHTOOL_OP_NEEDS_RTNL_TEST; } return false; } diff --git a/net/ipv4/arp.c b/net/ipv4/arp.c index d409f606aec0..60009d92e071 100644 --- a/net/ipv4/arp.c +++ b/net/ipv4/arp.c @@ -1278,6 +1278,7 @@ int arp_ioctl(struct net *net, unsigned int cmd, void __user *arg) err = copy_from_user(&r, arg, sizeof(struct arpreq)); if (err) return -EFAULT; + r.arp_dev[IFNAMSIZ - 1] = '\0'; break; default: return -EINVAL; diff --git a/net/ipv4/devinet.c b/net/ipv4/devinet.c index a35b72662e43..a80896647154 100644 --- a/net/ipv4/devinet.c +++ b/net/ipv4/devinet.c @@ -2117,9 +2117,10 @@ static int inet_validate_link_af(const struct net_device *dev, return err; if (tb[IFLA_INET_CONF]) { - err = nla_parse_nested(nested_tb, IPV4_DEVCONF_MAX, - tb[IFLA_INET_CONF], inet_devconf_policy, - extack); + err = nla_parse(nested_tb, IPV4_DEVCONF_MAX, + nla_data(tb[IFLA_INET_CONF]), + nla_len(tb[IFLA_INET_CONF]), + inet_devconf_policy, extack); if (err < 0) return err; diff --git a/net/ipv4/fib_semantics.c b/net/ipv4/fib_semantics.c index 0483519b7fb0..0c25f6dfb8a6 100644 --- a/net/ipv4/fib_semantics.c +++ b/net/ipv4/fib_semantics.c @@ -2176,6 +2176,15 @@ static bool fib_good_nh(const struct fib_nh *nh) return !!(state & NUD_VALID); } +static __be32 fib_nh_saddr(struct net *net, const struct fib_info *fi, + struct fib_nh *nh, int genid) +{ + if (READ_ONCE(nh->nh_saddr_genid) == genid) + return READ_ONCE(nh->nh_saddr); + + return fib_info_update_nhc_saddr(net, &nh->nh_common, fi->fib_scope); +} + void fib_select_multipath(struct fib_result *res, int hash, const struct flowi4 *fl4) { @@ -2184,6 +2193,7 @@ void fib_select_multipath(struct fib_result *res, int hash, bool use_neigh; int score = -1; __be32 saddr; + int genid; if (unlikely(res->fi->nh)) { nexthop_path_fib_result(res, hash); @@ -2192,6 +2202,7 @@ void fib_select_multipath(struct fib_result *res, int hash, use_neigh = READ_ONCE(net->ipv4.sysctl_fib_multipath_use_neigh); saddr = fl4 ? fl4->saddr : 0; + genid = saddr ? atomic_read(&net->ipv4.dev_addr_genid) : 0; change_nexthops(fi) { int nh_upper_bound, nh_score = 0; @@ -2204,7 +2215,7 @@ void fib_select_multipath(struct fib_result *res, int hash, (use_neigh && !fib_good_nh(nexthop_nh))) continue; - if (saddr && nexthop_nh->nh_saddr == saddr) + if (saddr && fib_nh_saddr(net, fi, nexthop_nh, genid) == saddr) nh_score += 2; if (hash <= nh_upper_bound) nh_score++; diff --git a/net/ipv4/fou_core.c b/net/ipv4/fou_core.c index ab09dfcdecbd..9a2999285b4a 100644 --- a/net/ipv4/fou_core.c +++ b/net/ipv4/fou_core.c @@ -600,6 +600,10 @@ static int fou_create(struct net *net, struct fou_cfg *cfg, /* Initial for fou type */ switch (cfg->type) { case FOU_ENCAP_DIRECT: + if (!cfg->protocol) { + err = -EINVAL; + goto error; + } tunnel_cfg.encap_rcv = fou_udp_recv; tunnel_cfg.gro_receive = fou_gro_receive; tunnel_cfg.gro_complete = fou_gro_complete; diff --git a/net/ipv4/inet_connection_sock.c b/net/ipv4/inet_connection_sock.c index 6257459bcee2..6a30f1138454 100644 --- a/net/ipv4/inet_connection_sock.c +++ b/net/ipv4/inet_connection_sock.c @@ -1520,7 +1520,8 @@ void inet_csk_listen_stop(struct sock *sk) local_bh_enable(); sock_put(child); - cond_resched(); + if (!has_current_bpf_ctx()) + cond_resched(); } if (queue->fastopenq.rskq_rst_head) { /* Free all the reqs queued in rskq_rst_head. */ diff --git a/net/ipv4/ip_gre.c b/net/ipv4/ip_gre.c index 0ba1e94e9012..384c02b9700b 100644 --- a/net/ipv4/ip_gre.c +++ b/net/ipv4/ip_gre.c @@ -1460,6 +1460,12 @@ static int ipgre_changelink(struct net_device *dev, struct nlattr *tb[], if (!rtnl_dev_link_net_capable(dev, t->net)) return -EPERM; + if (data && data[IFLA_GRE_COLLECT_METADATA] && !t->collect_md) { + NL_SET_ERR_MSG(extack, + "Enabling collect_md on an existing device is not supported"); + return -EOPNOTSUPP; + } + err = ipgre_newlink_encap_setup(dev, data); if (err) return err; @@ -1492,6 +1498,12 @@ static int erspan_changelink(struct net_device *dev, struct nlattr *tb[], if (!rtnl_dev_link_net_capable(dev, t->net)) return -EPERM; + if (data && data[IFLA_GRE_COLLECT_METADATA] && !t->collect_md) { + NL_SET_ERR_MSG(extack, + "Enabling collect_md on an existing device is not supported"); + return -EOPNOTSUPP; + } + err = ipgre_newlink_encap_setup(dev, data); if (err) return err; diff --git a/net/ipv4/ipconfig.c b/net/ipv4/ipconfig.c index a35ffedacc7c..b8a0c98e5e55 100644 --- a/net/ipv4/ipconfig.c +++ b/net/ipv4/ipconfig.c @@ -676,6 +676,24 @@ static const u8 ic_bootp_cookie[4] = { 99, 130, 83, 99 }; #ifdef IPCONFIG_DHCP +static bool __init +ic_dhcp_add_option(u8 **options, const u8 *end, u8 type, const void *value, + int len) +{ + u8 *e = *options; + + /* leave room for the option header and the END marker */ + if (len > U8_MAX || end - e < len + 3) + return false; + + *e++ = type; + *e++ = len; + memcpy(e, value, len); + *options = e + len; + + return true; +} + static void __init ic_dhcp_init_options(u8 *options, struct ic_device *d) { @@ -691,6 +709,7 @@ ic_dhcp_init_options(u8 *options, struct ic_device *d) 42, /* NTP servers */ }; u8 mt = (ic_servaddr == NONE) ? DHCPDISCOVER : DHCPREQUEST; + u8 *end = options + sizeof(((struct bootp_pkt *)0)->exten); u8 *e = options; int len; @@ -721,31 +740,19 @@ ic_dhcp_init_options(u8 *options, struct ic_device *d) e += sizeof(ic_req_params); if (ic_host_name_set) { - *e++ = 12; /* host-name */ len = strlen(utsname()->nodename); - *e++ = len; - memcpy(e, utsname()->nodename, len); - e += len; + ic_dhcp_add_option(&e, end, 12, utsname()->nodename, len); } if (*vendor_class_identifier) { - pr_info("DHCP: sending class identifier \"%s\"\n", - vendor_class_identifier); - *e++ = 60; /* Class-identifier */ len = strlen(vendor_class_identifier); - *e++ = len; - memcpy(e, vendor_class_identifier, len); - e += len; + if (ic_dhcp_add_option(&e, end, 60, vendor_class_identifier, len)) + pr_info("DHCP: sending class identifier \"%s\"\n", + vendor_class_identifier); } len = strlen(dhcp_client_identifier + 1); - /* the minimum length of identifier is 2, include 1 byte type, - * and can not be larger than the length of options - */ - if (len >= 1 && len < 312 - (e - options) - 1) { - *e++ = 61; - *e++ = len + 1; - memcpy(e, dhcp_client_identifier, len + 1); - e += len + 1; - } + /* the minimum length of identifier is 2, include 1 byte type */ + if (len >= 1) + ic_dhcp_add_option(&e, end, 61, dhcp_client_identifier, len + 1); *e++ = 255; /* End of the list */ } diff --git a/net/ipv4/tcp_output.c b/net/ipv4/tcp_output.c index f72a6b1fe749..0b89b624af8e 100644 --- a/net/ipv4/tcp_output.c +++ b/net/ipv4/tcp_output.c @@ -3886,6 +3886,7 @@ void tcp_send_active_reset(struct sock *sk, enum sk_rst_reason reason) */ int tcp_send_synack(struct sock *sk) { + struct tcp_sock *tp = tcp_sk(sk); struct sk_buff *skb; skb = tcp_rtx_queue_head(sk); @@ -3903,6 +3904,8 @@ int tcp_send_synack(struct sock *sk) if (!nskb) return -ENOMEM; INIT_LIST_HEAD(&nskb->tcp_tsorted_anchor); + if (skb == tp->retransmit_skb_hint) + tp->retransmit_skb_hint = nskb; tcp_highest_sack_replace(sk, skb, nskb); tcp_rtx_queue_unlink_and_free(skb, sk); __skb_header_release(nskb); diff --git a/net/ipv4/udp.c b/net/ipv4/udp.c index 45e96b721989..59afce196079 100644 --- a/net/ipv4/udp.c +++ b/net/ipv4/udp.c @@ -616,14 +616,23 @@ void udp_lib_hash4(struct sock *sk, u16 hash) struct net *net = sock_net(sk); struct udp_table *udptable; - /* Connected udp socket can re-connect to another remote address, which - * will be handled by rehash. Thus no need to redo hash4 here. + udptable = net->ipv4.udp_table; + hslot = udp_hashslot(udptable, net, udp_sk(sk)->udp_port_hash); + + /* A connected socket can re-connect to another address. rehash() + * relocates it, but only runs when the local address changes, so a + * socket bound to a specific address would stay filed under the + * previous peer's hash. Move it here. */ - if (udp_hashed4(sk)) + if (udp_hashed4(sk)) { + if (udp_sk(sk)->udp_lrpa_hash != hash) { + spin_lock_bh(&hslot->lock); + udp_rehash4(udptable, sk, hash); + spin_unlock_bh(&hslot->lock); + } return; + } - udptable = net->ipv4.udp_table; - hslot = udp_hashslot(udptable, net, udp_sk(sk)->udp_port_hash); hslot2 = udp_hashslot2(udptable, udp_sk(sk)->udp_portaddr_hash); hslot4 = udp_hashslot4(udptable, hash); udp_sk(sk)->udp_lrpa_hash = hash; @@ -2184,9 +2193,31 @@ int __udp_disconnect(struct sock *sk, int flags) } EXPORT_SYMBOL(__udp_disconnect); +/* __udp_disconnect() takes a socket out of the 4-tuple hash table only via + * ->rehash() or ->unhash(), and neither runs for a socket bound to a + * specific address and port. Remove it here, before its peer is cleared. + */ +static void udp_unhash4_on_disconnect(struct sock *sk) +{ + struct net *net = sock_net(sk); + struct udp_table *udptable; + struct udp_hslot *hslot; + + if (!udp_hashed4(sk)) + return; + + udptable = net->ipv4.udp_table; + hslot = udp_hashslot(udptable, net, udp_sk(sk)->udp_port_hash); + + spin_lock_bh(&hslot->lock); + udp_unhash4(udptable, sk); + spin_unlock_bh(&hslot->lock); +} + int udp_disconnect(struct sock *sk, int flags) { lock_sock(sk); + udp_unhash4_on_disconnect(sk); __udp_disconnect(sk, flags); release_sock(sk); return 0; @@ -3098,6 +3129,7 @@ int udp_abort(struct sock *sk, int err) sk->sk_err = err; sk_error_report(sk); + udp_unhash4_on_disconnect(sk); __udp_disconnect(sk, 0); out: diff --git a/net/ipv6/exthdrs_core.c b/net/ipv6/exthdrs_core.c index 9d06d487e8b1..4a9748338cf4 100644 --- a/net/ipv6/exthdrs_core.c +++ b/net/ipv6/exthdrs_core.c @@ -278,6 +278,9 @@ int ipv6_find_hdr(const struct sk_buff *skb, unsigned int *offset, hdrlen = ipv6_optlen(hp); if (!found) { + if (skb->len - start < hdrlen) + return -EBADMSG; + nexthdr = hp->nexthdr; start += hdrlen; } diff --git a/net/ipv6/ip6_fib.c b/net/ipv6/ip6_fib.c index 7c5daea3f096..fc3da984a9ab 100644 --- a/net/ipv6/ip6_fib.c +++ b/net/ipv6/ip6_fib.c @@ -1043,8 +1043,8 @@ static void fib6_purge_rt(struct fib6_info *rt, struct fib6_node *fn, struct fib6_table *table = rt->fib6_table; /* Flush all cached dst in exception table */ - rt6_flush_exceptions(rt); fib6_drop_pcpu_from(rt); + rt6_flush_exceptions(rt); if (rt->nh) { spin_lock(&rt->nh->lock); diff --git a/net/ipv6/ip6_gre.c b/net/ipv6/ip6_gre.c index 678678fcb5da..e6913a7f6881 100644 --- a/net/ipv6/ip6_gre.c +++ b/net/ipv6/ip6_gre.c @@ -2280,7 +2280,7 @@ static int ip6erspan_changelink(struct net_device *dev, struct nlattr *tb[], return PTR_ERR(t); ip6erspan_set_version(data, &p); - ip6gre_tunnel_unlink_md(ign, t); + ip6erspan_tunnel_unlink_md(ign, t); ip6gre_tunnel_unlink(ign, t); ip6erspan_tnl_change(t, &p, !tb[IFLA_MTU]); ip6erspan_tunnel_link_md(ign, t); diff --git a/net/ipv6/netfilter/ip6t_rpfilter.c b/net/ipv6/netfilter/ip6t_rpfilter.c index 67c87a88cde4..b5def30c3127 100644 --- a/net/ipv6/netfilter/ip6t_rpfilter.c +++ b/net/ipv6/netfilter/ip6t_rpfilter.c @@ -61,7 +61,7 @@ static bool rpfilter_lookup_reverse6(struct net *net, const struct sk_buff *skb, fl6.flowi6_oif = dev->ifindex; rt = (void *)ip6_route_lookup(net, &fl6, skb, lookup_flags); - if (rt->dst.error) + if (rt->dst.error || !rt->rt6i_idev) goto out; if (rt->rt6i_flags & (RTF_REJECT|RTF_ANYCAST)) diff --git a/net/ipv6/netfilter/ip6t_rt.c b/net/ipv6/netfilter/ip6t_rt.c index 278b52752f36..2f6261d06a07 100644 --- a/net/ipv6/netfilter/ip6t_rt.c +++ b/net/ipv6/netfilter/ip6t_rt.c @@ -96,7 +96,8 @@ static bool rt_mt6(const struct sk_buff *skb, struct xt_action_param *par) unsigned int i = 0; for (temp = 0; - temp < (unsigned int)((hdrlen - 8) / 16); + temp < (unsigned int)((hdrlen - 8) / 16) && + i < rtinfo->addrnr; temp++) { ap = skb_header_pointer(skb, ptr @@ -112,8 +113,6 @@ static bool rt_mt6(const struct sk_buff *skb, struct xt_action_param *par) if (ipv6_addr_equal(ap, &rtinfo->addrs[i])) i++; - if (i == rtinfo->addrnr) - break; } if (i == rtinfo->addrnr) return ret; @@ -162,6 +161,12 @@ static int rt_mt6_check(const struct xt_mtchk_param *par) pr_debug("too many addresses specified\n"); return -EINVAL; } + + if ((rtinfo->flags & IP6T_RT_FST_MASK) && !rtinfo->addrnr) { + pr_info_ratelimited("address list match requested but addrnr is 0\n"); + return -EINVAL; + } + if ((rtinfo->flags & (IP6T_RT_RES | IP6T_RT_FST_MASK)) && (!(rtinfo->flags & IP6T_RT_TYP) || (rtinfo->rt_type != 0) || diff --git a/net/ipv6/route.c b/net/ipv6/route.c index ee707e48f4ef..8fddcc631fcb 100644 --- a/net/ipv6/route.c +++ b/net/ipv6/route.c @@ -139,6 +139,7 @@ void rt6_uncached_list_add(struct rt6_info *rt) { struct uncached_list *ul = raw_cpu_ptr(&rt6_uncached_list); + /* Set once and never cleared: non-NULL marks an uncached route. */ rt->dst.rt_uncached_list = ul; spin_lock_bh(&ul->lock); @@ -1729,6 +1730,11 @@ static int rt6_insert_exception(struct rt6_info *nrt, spin_lock_bh(&rt6_exception_lock); + if (f6i->fib6_destroying) { + err = -ENOENT; + goto out; + } + bucket = rcu_dereference_protected(nh->rt6i_exception_bucket, lockdep_is_held(&rt6_exception_lock)); if (!bucket) { @@ -2721,8 +2727,8 @@ struct dst_entry *ip6_route_output_flags(struct net *net, rcu_read_lock(); dst = ip6_route_output_flags_noref(net, sk, fl6, flags); rt6 = dst_rt6_info(dst); - /* For dst cached in uncached_list, refcnt is already taken. */ - if (list_empty(&rt6->dst.rt_uncached) && !dst_hold_safe(dst)) { + /* For an uncached dst, refcnt is already taken. */ + if (!rt6->dst.rt_uncached_list && !dst_hold_safe(dst)) { dst = &net->ipv6.ip6_null_entry->dst; dst_hold(dst); } @@ -2831,7 +2837,7 @@ INDIRECT_CALLABLE_SCOPE struct dst_entry *ip6_dst_check(struct dst_entry *dst, from = rcu_dereference(rt->from); if (from && (rt->rt6i_flags & RTF_PCPU || - unlikely(!list_empty(&rt->dst.rt_uncached)))) + unlikely(rt->dst.rt_uncached_list))) dst_ret = rt6_dst_from_check(rt, from, cookie); else dst_ret = rt6_check(rt, from, cookie); diff --git a/net/ipv6/seg6.c b/net/ipv6/seg6.c index 62a7eb779202..8c2b156c227a 100644 --- a/net/ipv6/seg6.c +++ b/net/ipv6/seg6.c @@ -138,8 +138,8 @@ void seg6_icmp_srh(struct sk_buff *skb, struct inet6_skb_parm *opt) static struct genl_family seg6_genl_family; static const struct nla_policy seg6_genl_policy[SEG6_ATTR_MAX + 1] = { - [SEG6_ATTR_DST] = { .type = NLA_BINARY, - .len = sizeof(struct in6_addr) }, + [SEG6_ATTR_DST] = + NLA_POLICY_EXACT_LEN(sizeof(struct in6_addr)), [SEG6_ATTR_DSTLEN] = { .type = NLA_S32, }, [SEG6_ATTR_HMACKEYID] = { .type = NLA_U32, }, [SEG6_ATTR_SECRET] = { .type = NLA_BINARY, }, diff --git a/net/llc/llc_c_ac.c b/net/llc/llc_c_ac.c index 724ecd741d4c..1aa7fe28acdd 100644 --- a/net/llc/llc_c_ac.c +++ b/net/llc/llc_c_ac.c @@ -437,7 +437,7 @@ int llc_conn_ac_resend_i_xxx_x_set_0_or_send_rr(struct sock *sk, if (likely(!rc)) llc_conn_send_pdu(sk, nskb); else - kfree_skb(skb); + kfree_skb(nskb); } if (rc) { nr = LLC_I_GET_NR(pdu); diff --git a/net/llc/llc_s_ac.c b/net/llc/llc_s_ac.c index 98deee560373..831998211b52 100644 --- a/net/llc/llc_s_ac.c +++ b/net/llc/llc_s_ac.c @@ -121,6 +121,8 @@ int llc_sap_action_send_xid_r(struct llc_sap *sap, struct sk_buff *skb) rc = llc_mac_hdr_init(nskb, mac_sa, mac_da); if (likely(!rc)) rc = dev_queue_xmit(nskb); + else + kfree_skb(nskb); out: return rc; } @@ -170,6 +172,8 @@ int llc_sap_action_send_test_r(struct llc_sap *sap, struct sk_buff *skb) rc = llc_mac_hdr_init(nskb, mac_sa, mac_da); if (likely(!rc)) rc = dev_queue_xmit(nskb); + else + kfree_skb(nskb); out: return rc; } diff --git a/net/llc/llc_sap.c b/net/llc/llc_sap.c index 1bd446a21092..3904a1b4ba84 100644 --- a/net/llc/llc_sap.c +++ b/net/llc/llc_sap.c @@ -19,12 +19,12 @@ #include #include -static int llc_mac_header_len(unsigned short devtype) +static int llc_mac_header_len(struct net_device *dev) { - switch (devtype) { + switch (dev->type) { case ARPHRD_ETHER: case ARPHRD_LOOPBACK: - return sizeof(struct ethhdr); + return LL_RESERVED_SPACE(dev); } return 0; } @@ -45,7 +45,7 @@ struct sk_buff *llc_alloc_frame(struct sock *sk, struct net_device *dev, int hlen = type == LLC_PDU_TYPE_U ? 3 : 4; struct sk_buff *skb; - hlen += llc_mac_header_len(dev->type); + hlen += llc_mac_header_len(dev); skb = alloc_skb(hlen + data_size, GFP_ATOMIC); if (skb) { diff --git a/net/mctp/route.c b/net/mctp/route.c index b19c63a5691a..e7c95eeacb48 100644 --- a/net/mctp/route.c +++ b/net/mctp/route.c @@ -825,7 +825,7 @@ static struct mctp_sk_key *mctp_lookup_prealloc_tag(struct mctp_sock *msk, spin_lock_irqsave(&mns->keys_lock, flags); - hlist_for_each_entry(tmp, &mns->keys, hlist) { + hlist_for_each_entry(tmp, &msk->keys, sklist) { if (tmp->net != netid) continue; diff --git a/net/mptcp/protocol.c b/net/mptcp/protocol.c index 49c46498af96..de6f289a6e14 100644 --- a/net/mptcp/protocol.c +++ b/net/mptcp/protocol.c @@ -2401,7 +2401,7 @@ static int mptcp_recvmsg(struct sock *sk, struct msghdr *msg, size_t len, mptcp_cleanup_rbuf(msk, copied); err = sk_wait_data(sk, &timeo, last); if (err < 0) { - err = copied ? : err; + copied = copied ? : err; goto out_err; } } diff --git a/net/netfilter/ipvs/ip_vs_core.c b/net/netfilter/ipvs/ip_vs_core.c index eb806813292a..6802ce5cb1de 100644 --- a/net/netfilter/ipvs/ip_vs_core.c +++ b/net/netfilter/ipvs/ip_vs_core.c @@ -1961,6 +1961,12 @@ ip_vs_in_icmp(struct netns_ipvs *ipvs, struct sk_buff *skb, int *related, /* Ensure the IP header is present in headroom */ if (!pskb_may_pull(skb, hlen_orig)) goto ignore_tunnel; + skb_set_transport_header(skb, hlen_orig); + /* Before now we may used ihl from skb frag, revalidate it after + * copying it into skb head to prevent out-of-bounds access + */ + if (ip_hdr(skb)->ihl * 4 != hlen_orig) + goto ignore_tunnel; IP_VS_DBG(12, "Sending ICMP for %pI4->%pI4: t=%u, c=%u, i=%u\n", &ip_hdr(skb)->saddr, &ip_hdr(skb)->daddr, type, code, ntohl(info)); diff --git a/net/netfilter/nf_conntrack_netlink.c b/net/netfilter/nf_conntrack_netlink.c index 92c3bb77d27e..6cddfa339d71 100644 --- a/net/netfilter/nf_conntrack_netlink.c +++ b/net/netfilter/nf_conntrack_netlink.c @@ -3392,7 +3392,8 @@ static bool expect_iter_name(struct nf_conntrack_expect *exp, void *data) struct nf_conntrack_helper *helper; const char *name = data; - helper = rcu_dereference(exp->helper); + helper = rcu_dereference_protected(exp->helper, + lockdep_is_held(&nf_conntrack_expect_lock)); if (!helper) return false; diff --git a/net/netfilter/nf_flow_table_offload.c b/net/netfilter/nf_flow_table_offload.c index 801a3dd9ceea..6757fd89c1f1 100644 --- a/net/netfilter/nf_flow_table_offload.c +++ b/net/netfilter/nf_flow_table_offload.c @@ -995,7 +995,6 @@ static void flow_offload_work_del(struct flow_offload_work *offload) flow_offload_tuple_del(offload, FLOW_OFFLOAD_DIR_ORIGINAL); if (test_bit(NF_FLOW_HW_BIDIRECTIONAL, &offload->flow->flags)) flow_offload_tuple_del(offload, FLOW_OFFLOAD_DIR_REPLY); - set_bit(NF_FLOW_HW_DEAD, &offload->flow->flags); } static void flow_offload_tuple_stats(struct flow_offload_work *offload, @@ -1059,6 +1058,12 @@ static void flow_offload_work_handler(struct work_struct *work) } clear_bit(NF_FLOW_HW_PENDING, &offload->flow->flags); + if (offload->cmd == FLOW_CLS_DESTROY) { + /* Publish after the worker's last flow access. */ + smp_mb__before_atomic(); + set_bit(NF_FLOW_HW_DEAD, &offload->flow->flags); + } + kfree(offload); } diff --git a/net/netfilter/nf_tables_api.c b/net/netfilter/nf_tables_api.c index 245fe7b2f673..ea15ebd7bce8 100644 --- a/net/netfilter/nf_tables_api.c +++ b/net/netfilter/nf_tables_api.c @@ -6990,11 +6990,14 @@ static int nft_setelem_catchall_insert(const struct net *net, { struct nft_set_elem_catchall *catchall; u8 genmask = nft_genmask_next(net); + u64 tstamp = nft_net_tstamp(net); struct nft_set_ext *ext; list_for_each_entry(catchall, &set->catchall_list, list) { ext = nft_set_elem_ext(set, catchall->elem); - if (nft_set_elem_active(ext, genmask)) { + if (nft_set_elem_active(ext, genmask) && + !__nft_set_elem_expired(ext, tstamp) && + !nft_set_elem_is_dead(ext)) { *priv = catchall->elem; return -EEXIST; } @@ -7087,11 +7090,14 @@ static int nft_setelem_catchall_deactivate(const struct net *net, struct nft_set_elem *elem) { struct nft_set_elem_catchall *catchall; + u64 tstamp = nft_net_tstamp(net); struct nft_set_ext *ext; list_for_each_entry(catchall, &set->catchall_list, list) { ext = nft_set_elem_ext(set, catchall->elem); - if (!nft_is_active_next(net, ext)) + if (!nft_is_active_next(net, ext) || + __nft_set_elem_expired(ext, tstamp) || + nft_set_elem_is_dead(ext)) continue; kfree(elem->priv); diff --git a/net/netfilter/nfnetlink_queue.c b/net/netfilter/nfnetlink_queue.c index b8aaf39cb4d8..ec5e7716df8b 100644 --- a/net/netfilter/nfnetlink_queue.c +++ b/net/netfilter/nfnetlink_queue.c @@ -1527,6 +1527,7 @@ nfqnl_rcv_nl_event(struct notifier_block *this, if (event == NETLINK_URELEASE && n->protocol == NETLINK_NETFILTER) { int i; + nfnl_lock(NFNL_SUBSYS_QUEUE); /* destroy all instances for this portid */ spin_lock(&q->instances_lock); for (i = 0; i < INSTANCE_BUCKETS; i++) { @@ -1540,6 +1541,7 @@ nfqnl_rcv_nl_event(struct notifier_block *this, } } spin_unlock(&q->instances_lock); + nfnl_unlock(NFNL_SUBSYS_QUEUE); } return NOTIFY_DONE; } @@ -1859,9 +1861,9 @@ static int nfqnl_recv_config(struct sk_buff *skb, const struct nfnl_info *info, /* Lookup queue under RCU. After peer_portid check (or for new queue * in BIND case), the queue is owned by the socket sending this message. - * A socket cannot simultaneously send a message and close, so while - * processing this CONFIG message, nfqnl_rcv_nl_event() (triggered by - * socket close) cannot destroy this queue. Safe to use without RCU. + * nfqnl_rcv_nl_event() will block on the nfnl subsys mutex that is + * held by the caller, so the queue cannot be destroyed in parallel, + * even after we drop the RCU read lock. */ rcu_read_lock(); queue = instance_lookup(q, queue_num); diff --git a/net/netfilter/nft_synproxy.c b/net/netfilter/nft_synproxy.c index 9ed288c9d168..554a96a000f4 100644 --- a/net/netfilter/nft_synproxy.c +++ b/net/netfilter/nft_synproxy.c @@ -118,7 +118,8 @@ static void nft_synproxy_do_eval(const struct nft_synproxy *priv, return; } - if (nf_ip_checksum(skb, nft_hook(pkt), thoff, IPPROTO_TCP)) { + if (nf_checksum(skb, nft_hook(pkt), thoff, IPPROTO_TCP, + nft_pf(pkt))) { regs->verdict.code = NF_DROP; return; } diff --git a/net/netlink/genetlink.c b/net/netlink/genetlink.c index 41d37442f186..5cc1037d4917 100644 --- a/net/netlink/genetlink.c +++ b/net/netlink/genetlink.c @@ -1656,7 +1656,7 @@ static void *ctrl_dumppolicy_prep(struct sk_buff *skb, } static int ctrl_dumppolicy_put_op(struct sk_buff *skb, - struct netlink_callback *cb, + struct netlink_callback *cb, u32 cmd, struct genl_split_ops *doit, struct genl_split_ops *dumpit) { @@ -1677,7 +1677,7 @@ static int ctrl_dumppolicy_put_op(struct sk_buff *skb, if (!nest_pol) goto err; - nest_op = nla_nest_start(skb, doit->cmd); + nest_op = nla_nest_start(skb, cmd); if (!nest_op) goto err; @@ -1721,7 +1721,8 @@ static int ctrl_dumppolicy(struct sk_buff *skb, struct netlink_callback *cb) &doit, &dumpit))) return -ENOENT; - if (ctrl_dumppolicy_put_op(skb, cb, &doit, &dumpit)) + if (ctrl_dumppolicy_put_op(skb, cb, ctx->op, + &doit, &dumpit)) return skb->len; /* done with the per-op policy index list */ @@ -1730,6 +1731,7 @@ static int ctrl_dumppolicy(struct sk_buff *skb, struct netlink_callback *cb) while (ctx->dump_map) { if (ctrl_dumppolicy_put_op(skb, cb, + ctx->op_iter->cmd, &ctx->op_iter->doit, &ctx->op_iter->dumpit)) return skb->len; diff --git a/net/nfc/core.c b/net/nfc/core.c index a92a6566e6a0..f521669293f0 100644 --- a/net/nfc/core.c +++ b/net/nfc/core.c @@ -279,10 +279,10 @@ static struct nfc_target *nfc_find_target(struct nfc_dev *dev, u32 target_idx) int nfc_dep_link_up(struct nfc_dev *dev, int target_index, u8 comm_mode) { - int rc = 0; - u8 *gb; - size_t gb_len; struct nfc_target *target; + u8 gb[NFC_MAX_GT_LEN]; + size_t gb_len = 0; + int rc = 0; pr_debug("dev_name=%s comm %d\n", dev_name(&dev->dev), comm_mode); @@ -301,7 +301,7 @@ int nfc_dep_link_up(struct nfc_dev *dev, int target_index, u8 comm_mode) goto error; } - gb = nfc_llcp_general_bytes(dev, &gb_len); + nfc_get_local_general_bytes(dev, gb, sizeof(gb), &gb_len); if (gb_len > NFC_MAX_GT_LEN) { rc = -EINVAL; goto error; @@ -644,11 +644,10 @@ int nfc_set_remote_general_bytes(struct nfc_dev *dev, const u8 *gb, u8 gb_len) } EXPORT_SYMBOL(nfc_set_remote_general_bytes); -u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, size_t *gb_len) +u8 *nfc_get_local_general_bytes(struct nfc_dev *dev, u8 *out_gb, + size_t gb_max_len, size_t *gb_len) { - pr_debug("dev_name=%s\n", dev_name(&dev->dev)); - - return nfc_llcp_general_bytes(dev, gb_len); + return nfc_llcp_general_bytes(dev, out_gb, gb_max_len, gb_len); } EXPORT_SYMBOL(nfc_get_local_general_bytes); diff --git a/net/nfc/digital_dep.c b/net/nfc/digital_dep.c index 3982fa084737..968547c306a5 100644 --- a/net/nfc/digital_dep.c +++ b/net/nfc/digital_dep.c @@ -1490,14 +1490,14 @@ static int digital_tg_send_atr_res(struct nfc_digital_dev *ddev, struct digital_atr_req *atr_req) { struct digital_atr_res *atr_res; + u8 gb[NFC_MAX_GT_LEN]; struct sk_buff *skb; - u8 *gb, payload_bits; + u8 payload_bits; size_t gb_len; int rc; - gb = nfc_get_local_general_bytes(ddev->nfc_dev, &gb_len); - if (!gb) - gb_len = 0; + nfc_get_local_general_bytes(ddev->nfc_dev, gb, sizeof(gb), + &gb_len); skb = digital_skb_alloc(ddev, sizeof(struct digital_atr_res) + gb_len); if (!skb) diff --git a/net/nfc/llcp.h b/net/nfc/llcp.h index d8345ed57c95..23ae7a0112d3 100644 --- a/net/nfc/llcp.h +++ b/net/nfc/llcp.h @@ -91,6 +91,7 @@ struct nfc_llcp_local { struct hlist_head pending_sdreqs; struct timer_list sdreq_timer; struct work_struct sdreq_timeout_work; + struct work_struct release_work; u8 sdreq_next_tid; /* sockets array */ diff --git a/net/nfc/llcp_commands.c b/net/nfc/llcp_commands.c index ca89fe967d6a..80a00938c869 100644 --- a/net/nfc/llcp_commands.c +++ b/net/nfc/llcp_commands.c @@ -135,7 +135,7 @@ struct nfc_llcp_sdp_tlv *nfc_llcp_build_sdreq_tlv(u8 tid, const char *uri, { struct nfc_llcp_sdp_tlv *sdreq; - pr_debug("uri: %s, len: %zu\n", uri, uri_len); + pr_debug("uri: %.*s, len: %zu\n", (int)uri_len, uri, uri_len); /* sdreq->tlv_len is u8, takes uri_len, + 3 for header, + 1 for NULL */ if (WARN_ON_ONCE(uri_len > U8_MAX - 4)) diff --git a/net/nfc/llcp_core.c b/net/nfc/llcp_core.c index cac1b5487064..74bf817007cf 100644 --- a/net/nfc/llcp_core.c +++ b/net/nfc/llcp_core.c @@ -20,6 +20,8 @@ static LIST_HEAD(llcp_devices); /* Protects llcp_devices list */ static DEFINE_SPINLOCK(llcp_devices_lock); +static struct workqueue_struct *llcp_wq; + static void nfc_llcp_rx_skb(struct nfc_llcp_local *local, struct sk_buff *skb); void nfc_llcp_sock_link(struct llcp_sock_list *l, struct sock *sk) @@ -63,21 +65,33 @@ static void nfc_llcp_socket_purge(struct nfc_llcp_sock *sock) } } +static struct sock *nfc_llcp_sock_list_pop(struct llcp_sock_list *l) +{ + struct sock *sk; + + write_lock(&l->lock); + sk = sk_head(&l->head); + if (sk) { + sock_hold(sk); + sk_del_node_init(sk); + } + write_unlock(&l->lock); + + return sk; +} + static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, int err) { struct sock *sk; - struct hlist_node *tmp; struct nfc_llcp_sock *llcp_sock; skb_queue_purge(&local->tx_queue); - write_lock(&local->sockets.lock); - - sk_for_each_safe(sk, tmp, &local->sockets.head) { + while ((sk = nfc_llcp_sock_list_pop(&local->sockets))) { llcp_sock = nfc_llcp_sock(sk); - bh_lock_sock(sk); + lock_sock(sk); nfc_llcp_socket_purge(llcp_sock); @@ -91,17 +105,27 @@ static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, list_for_each_entry_safe(lsk, n, &llcp_sock->accept_queue, accept_queue) { - accept_sk = &lsk->sk; - bh_lock_sock(accept_sk); + bool put_creation = false; - nfc_llcp_accept_unlink(accept_sk); - - if (err) - accept_sk->sk_err = err; - accept_sk->sk_state = LLCP_CLOSED; - accept_sk->sk_state_change(sk); + accept_sk = &lsk->sk; + lock_sock_nested(accept_sk, + SINGLE_DEPTH_NESTING); + + if (nfc_llcp_sock(accept_sk)->parent == sk) { + nfc_llcp_accept_unlink(accept_sk); + nfc_llcp_sock_unlink(&local->sockets, accept_sk); + + if (err) + accept_sk->sk_err = err; + accept_sk->sk_state = LLCP_CLOSED; + accept_sk->sk_state_change(accept_sk); + sock_orphan(accept_sk); + put_creation = true; + } - bh_unlock_sock(accept_sk); + release_sock(accept_sk); + if (put_creation) + sock_put(accept_sk); /* creation ref */ } } @@ -110,23 +134,18 @@ static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, sk->sk_state = LLCP_CLOSED; sk->sk_state_change(sk); - bh_unlock_sock(sk); - - sk_del_node_init(sk); + release_sock(sk); + sock_put(sk); } - write_unlock(&local->sockets.lock); - /* If we still have a device, we keep the RAW sockets alive */ if (device == true) return; - write_lock(&local->raw_sockets.lock); - - sk_for_each_safe(sk, tmp, &local->raw_sockets.head) { + while ((sk = nfc_llcp_sock_list_pop(&local->raw_sockets))) { llcp_sock = nfc_llcp_sock(sk); - bh_lock_sock(sk); + lock_sock(sk); nfc_llcp_socket_purge(llcp_sock); @@ -135,26 +154,20 @@ static void nfc_llcp_socket_release(struct nfc_llcp_local *local, bool device, sk->sk_state = LLCP_CLOSED; sk->sk_state_change(sk); - bh_unlock_sock(sk); - - sk_del_node_init(sk); + release_sock(sk); + sock_put(sk); } - - write_unlock(&local->raw_sockets.lock); } static struct nfc_llcp_local *nfc_llcp_local_get(struct nfc_llcp_local *local) { - /* Since using nfc_llcp_local may result in usage of nfc_dev, whenever - * we hold a reference to local, we also need to hold a reference to - * the device to avoid UAF. - */ - if (!nfc_get_device(local->dev->idx)) + if (!local) return NULL; - kref_get(&local->ref); + if (kref_get_unless_zero(&local->ref)) + return local; - return local; + return NULL; } static void local_cleanup(struct nfc_llcp_local *local) @@ -172,30 +185,34 @@ static void local_cleanup(struct nfc_llcp_local *local) nfc_llcp_free_sdp_tlv_list(&local->pending_sdreqs); } -static void local_release(struct kref *ref) +static void local_release_work(struct work_struct *work) { struct nfc_llcp_local *local; + struct nfc_dev *dev; - local = container_of(ref, struct nfc_llcp_local, ref); + local = container_of(work, struct nfc_llcp_local, release_work); + dev = local->dev; local_cleanup(local); kfree(local); + nfc_put_device(dev); } -int nfc_llcp_local_put(struct nfc_llcp_local *local) +static void local_release(struct kref *ref) { - struct nfc_dev *dev; - int ret; + struct nfc_llcp_local *local; - if (local == NULL) - return 0; + local = container_of(ref, struct nfc_llcp_local, ref); - dev = local->dev; + queue_work(llcp_wq, &local->release_work); +} - ret = kref_put(&local->ref, local_release); - nfc_put_device(dev); +int nfc_llcp_local_put(struct nfc_llcp_local *local) +{ + if (!local) + return 0; - return ret; + return kref_put(&local->ref, local_release); } static struct nfc_llcp_sock *nfc_llcp_sock_get(struct nfc_llcp_local *local, @@ -341,7 +358,7 @@ static int nfc_llcp_wks_sap(const char *service_name, size_t service_name_len) { int sap, num_wks; - pr_debug("%s\n", service_name); + pr_debug("%.*s\n", (int)service_name_len, service_name); if (service_name == NULL) return -EINVAL; @@ -352,7 +369,8 @@ static int nfc_llcp_wks_sap(const char *service_name, size_t service_name_len) if (wks[sap] == NULL) continue; - if (strncmp(wks[sap], service_name, service_name_len) == 0) + if (strlen(wks[sap]) == service_name_len && + !strncmp(wks[sap], service_name, service_name_len)) return sap; } @@ -635,23 +653,32 @@ static int nfc_llcp_build_gb(struct nfc_llcp_local *local) return ret; } -u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, size_t *general_bytes_len) +u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, u8 *out_gb, size_t gb_max_len, + size_t *general_bytes_len) { struct nfc_llcp_local *local; + if (!out_gb || !general_bytes_len) + return NULL; + local = nfc_llcp_find_local(dev); - if (local == NULL) { + if (!local) { *general_bytes_len = 0; return NULL; } nfc_llcp_build_gb(local); - *general_bytes_len = local->gb_len; + if (local->gb_len) { + *general_bytes_len = min_t(size_t, local->gb_len, gb_max_len); + memcpy(out_gb, local->gb, *general_bytes_len); + } else { + *general_bytes_len = 0; + } nfc_llcp_local_put(local); - return local->gb; + return out_gb; } int nfc_llcp_set_remote_gb(struct nfc_dev *dev, const u8 *gb, u8 gb_len) @@ -1074,6 +1101,9 @@ static void nfc_llcp_recv_hdlc(struct nfc_llcp_local *local, struct sock *sk; u8 dsap, ssap, ptype, ns, nr; + if (!pskb_may_pull(skb, LLCP_HEADER_SIZE + LLCP_SEQUENCE_SIZE)) + return; + ptype = nfc_llcp_ptype(skb); dsap = nfc_llcp_dsap(skb); ssap = nfc_llcp_ssap(skb); @@ -1251,6 +1281,7 @@ static void nfc_llcp_recv_dm(struct nfc_llcp_local *local, struct nfc_llcp_sock *llcp_sock; struct sock *sk; u8 dsap, ssap, reason; + bool connecting = false; dsap = nfc_llcp_dsap(skb); ssap = nfc_llcp_ssap(skb); @@ -1262,6 +1293,7 @@ static void nfc_llcp_recv_dm(struct nfc_llcp_local *local, case LLCP_DM_NOBOUND: case LLCP_DM_REJ: llcp_sock = nfc_llcp_connecting_sock_get(local, dsap); + connecting = true; break; default: @@ -1276,10 +1308,33 @@ static void nfc_llcp_recv_dm(struct nfc_llcp_local *local, sk = &llcp_sock->sk; + lock_sock(sk); + + /* Check if socket was destroyed whilst waiting for the lock */ + if (!sk_hashed(sk)) { + release_sock(sk); + nfc_llcp_sock_put(llcp_sock); + return; + } + + /* + * For DM(NOBOUND)/DM(REJ) the socket is still linked on the + * connecting_sockets list. Unlink it here, under the socket lock, + * before moving it to LLCP_CLOSED: llcp_sock_release() selects the + * list to unlink from by sk_state, so leaving a connecting socket + * in the CLOSED state would make it unlink from the wrong list and + * corrupt the connecting_sockets list / desync the socket refcount. + * This mirrors nfc_llcp_recv_cc(). + */ + if (connecting) + nfc_llcp_sock_unlink(&local->connecting_sockets, sk); + sk->sk_err = ENXIO; sk->sk_state = LLCP_CLOSED; sk->sk_state_change(sk); + release_sock(sk); + nfc_llcp_sock_put(llcp_sock); } @@ -1680,6 +1735,7 @@ int nfc_llcp_register_device(struct nfc_dev *ndev) INIT_WORK(&local->rx_work, nfc_llcp_rx_work); INIT_WORK(&local->timeout_work, nfc_llcp_timeout_work); + INIT_WORK(&local->release_work, local_release_work); rwlock_init(&local->sockets.lock); rwlock_init(&local->connecting_sockets.lock); @@ -1723,10 +1779,23 @@ void nfc_llcp_unregister_device(struct nfc_dev *dev) int __init nfc_llcp_init(void) { - return nfc_llcp_sock_init(); + int ret; + + llcp_wq = alloc_workqueue("nfc_llcp_wq", WQ_UNBOUND, 0); + if (!llcp_wq) + return -ENOMEM; + + ret = nfc_llcp_sock_init(); + if (ret) { + destroy_workqueue(llcp_wq); + return ret; + } + + return 0; } void nfc_llcp_exit(void) { nfc_llcp_sock_exit(); + destroy_workqueue(llcp_wq); } diff --git a/net/nfc/llcp_sock.c b/net/nfc/llcp_sock.c index 5558d8a4d48b..1e5ee4bcde68 100644 --- a/net/nfc/llcp_sock.c +++ b/net/nfc/llcp_sock.c @@ -392,11 +392,12 @@ void nfc_llcp_accept_unlink(struct sock *sk) pr_debug("state %d\n", sk->sk_state); - list_del_init(&llcp_sock->accept_queue); - sk_acceptq_removed(llcp_sock->parent); - llcp_sock->parent = NULL; - - sock_put(sk); + if (llcp_sock->parent) { + list_del_init(&llcp_sock->accept_queue); + sk_acceptq_removed(llcp_sock->parent); + llcp_sock->parent = NULL; + sock_put(sk); + } } void nfc_llcp_accept_enqueue(struct sock *parent, struct sock *sk) @@ -423,12 +424,20 @@ struct sock *nfc_llcp_accept_dequeue(struct sock *parent, list_for_each_entry_safe(lsk, n, &llcp_parent->accept_queue, accept_queue) { + struct nfc_llcp_local *local; + sk = &lsk->sk; - lock_sock(sk); + lock_sock_nested(sk, SINGLE_DEPTH_NESTING); if (sk->sk_state == LLCP_CLOSED) { - release_sock(sk); + local = nfc_llcp_sock(sk)->local; + nfc_llcp_accept_unlink(sk); + if (local) + nfc_llcp_sock_unlink(&local->sockets, sk); + sock_orphan(sk); + release_sock(sk); + sock_put(sk); continue; } @@ -464,7 +473,7 @@ static int llcp_sock_accept(struct socket *sock, struct socket *newsock, pr_debug("parent %p\n", sk); - lock_sock_nested(sk, SINGLE_DEPTH_NESTING); + lock_sock(sk); if (sk->sk_state != LLCP_LISTEN) { ret = -EBADFD; @@ -490,7 +499,12 @@ static int llcp_sock_accept(struct socket *sock, struct socket *newsock, release_sock(sk); timeo = schedule_timeout(timeo); - lock_sock_nested(sk, SINGLE_DEPTH_NESTING); + lock_sock(sk); + + if (sk->sk_state != LLCP_LISTEN) { + ret = -EBADFD; + break; + } } __set_current_state(TASK_RUNNING); remove_wait_queue(sk_sleep(sk), &wait); @@ -629,13 +643,24 @@ static int llcp_sock_release(struct socket *sock) list_for_each_entry_safe(lsk, n, &llcp_sock->accept_queue, accept_queue) { + bool put_creation = false; + accept_sk = &lsk->sk; - lock_sock(accept_sk); + lock_sock_nested(accept_sk, SINGLE_DEPTH_NESTING); + + if (nfc_llcp_sock(accept_sk)->parent == sk) { + nfc_llcp_send_disconnect(lsk); + nfc_llcp_accept_unlink(accept_sk); + nfc_llcp_sock_unlink(&local->sockets, accept_sk); - nfc_llcp_send_disconnect(lsk); - nfc_llcp_accept_unlink(accept_sk); + accept_sk->sk_state = LLCP_CLOSED; + sock_orphan(accept_sk); + put_creation = true; + } release_sock(accept_sk); + if (put_creation) + sock_put(accept_sk); /* creation ref */ } } @@ -734,12 +759,16 @@ static int llcp_sock_connect(struct socket *sock, struct sockaddr_unsized *_addr llcp_sock->service_name_len = min_t(unsigned int, addr->service_name_len, NFC_LLCP_MAX_SERVICE_NAME); - llcp_sock->service_name = kmemdup(addr->service_name, - llcp_sock->service_name_len, - GFP_KERNEL); - if (!llcp_sock->service_name) { - ret = -ENOMEM; - goto sock_llcp_release; + if (llcp_sock->service_name_len == 0) { + llcp_sock->service_name = NULL; + } else { + llcp_sock->service_name = kmemdup(addr->service_name, + llcp_sock->service_name_len, + GFP_KERNEL); + if (!llcp_sock->service_name) { + ret = -ENOMEM; + goto sock_llcp_release; + } } nfc_llcp_sock_link(&local->connecting_sockets, sk); diff --git a/net/nfc/nci/core.c b/net/nfc/nci/core.c index 5f46c4b5720f..73e3a96470ac 100644 --- a/net/nfc/nci/core.c +++ b/net/nfc/nci/core.c @@ -780,15 +780,15 @@ static int nci_set_local_general_bytes(struct nfc_dev *nfc_dev) { struct nci_dev *ndev = nfc_get_drvdata(nfc_dev); struct nci_set_config_param param; + u8 gb[NFC_MAX_GT_LEN]; int rc; - param.val = nfc_get_local_general_bytes(nfc_dev, ¶m.len); - if ((param.val == NULL) || (param.len == 0)) + nfc_get_local_general_bytes(nfc_dev, gb, sizeof(gb), + ¶m.len); + if (param.len == 0) return 0; - if (param.len > NFC_MAX_GT_LEN) - return -EINVAL; - + param.val = gb; param.id = NCI_PN_ATR_REQ_GEN_BYTES; rc = nci_request(ndev, nci_set_config_req, ¶m, diff --git a/net/nfc/netlink.c b/net/nfc/netlink.c index 0c58824cb150..224bdfa2dd0d 100644 --- a/net/nfc/netlink.c +++ b/net/nfc/netlink.c @@ -1181,7 +1181,7 @@ static int nfc_genl_llc_sdreq(struct sk_buff *skb, struct genl_info *info) if (rc != 0) { rc = -EINVAL; - goto put_local; + goto free_list; } if (!sdp_attrs[NFC_SDP_ATTR_URI]) @@ -1200,7 +1200,7 @@ static int nfc_genl_llc_sdreq(struct sk_buff *skb, struct genl_info *info) sdreq = nfc_llcp_build_sdreq_tlv(tid, uri, uri_len); if (sdreq == NULL) { rc = -ENOMEM; - goto put_local; + goto free_list; } tlvs_len += sdreq->tlv_len; @@ -1215,6 +1215,9 @@ static int nfc_genl_llc_sdreq(struct sk_buff *skb, struct genl_info *info) rc = nfc_llcp_send_snl_sdreq(local, &sdreq_list, tlvs_len); +free_list: + nfc_llcp_free_sdp_tlv_list(&sdreq_list); + put_local: nfc_llcp_local_put(local); diff --git a/net/nfc/nfc.h b/net/nfc/nfc.h index 0b1e6466f4fb..82c5dfdad10e 100644 --- a/net/nfc/nfc.h +++ b/net/nfc/nfc.h @@ -49,7 +49,8 @@ void nfc_llcp_mac_is_up(struct nfc_dev *dev, u32 target_idx, int nfc_llcp_register_device(struct nfc_dev *dev); void nfc_llcp_unregister_device(struct nfc_dev *dev); int nfc_llcp_set_remote_gb(struct nfc_dev *dev, const u8 *gb, u8 gb_len); -u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, size_t *general_bytes_len); +u8 *nfc_llcp_general_bytes(struct nfc_dev *dev, u8 *out_gb, size_t gb_max_len, + size_t *general_bytes_len); int nfc_llcp_data_received(struct nfc_dev *dev, struct sk_buff *skb); struct nfc_llcp_local *nfc_llcp_find_local(struct nfc_dev *dev); int nfc_llcp_local_put(struct nfc_llcp_local *local); diff --git a/net/openvswitch/conntrack.c b/net/openvswitch/conntrack.c index 757cce7a658e..4afb750ad2c8 100644 --- a/net/openvswitch/conntrack.c +++ b/net/openvswitch/conntrack.c @@ -734,6 +734,18 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, enum ip_conntrack_info ctinfo; struct nf_conn *ct; + /* If the ct entry is not confirmed and shared with some other skb, + * e.g., a cloned one, we can't just modify it with the commit as we + * must not modify the extension set. Reset. + */ + if (cached && info->commit) { + ct = nf_ct_get(skb, &ctinfo); + if (ct && !nf_ct_is_confirmed(ct) && nf_ct_shared(ct)) { + nf_reset_ct(skb); + cached = false; + } + } + if (!cached) { struct nf_hook_state state = { .hook = NF_INET_PRE_ROUTING, @@ -767,8 +779,6 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, ct = nf_ct_get(skb, &ctinfo); if (ct) { - bool add_helper = false; - /* Packets starting a new connection must be NATted before the * helper, so that the helper knows about the NAT. We enforce * this by delaying both NAT and helper calls for unconfirmed @@ -800,7 +810,6 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, GFP_ATOMIC); if (err) return err; - add_helper = true; /* helper installed, add seqadj if NAT is required */ if (info->nat && !nfct_seqadj(ct)) { @@ -809,14 +818,14 @@ static int __ovs_ct_lookup(struct net *net, struct sw_flow_key *key, } } - /* Call the helper only if: - * - nf_conntrack_in() was executed above ("!cached") or a - * helper was just attached ("add_helper") for a confirmed - * connection, or - * - When committing an unconfirmed connection. + /* Call the helper only if nf_conntrack_in() was executed + * above ("!cached"). + * + * For unconfirmed connections it will be called later during + * commit as we need to have all the other extensions allocated + * before the call. */ - if ((nf_ct_is_confirmed(ct) ? !cached || add_helper : - info->commit)) { + if (nf_ct_is_confirmed(ct) && !cached) { int err = nf_ct_helper(skb, ct, ctinfo, info->family); err = verdict_to_errno(err); @@ -1020,6 +1029,14 @@ static int ovs_ct_commit(struct net *net, struct sw_flow_key *key, return err; nf_conn_act_ct_ext_add(skb, ct, ctinfo); + + /* Call the helpers now. We couldn't do this before as + * all the extensions must be allocated before the call. + */ + err = nf_ct_helper(skb, ct, ctinfo, info->family); + err = verdict_to_errno(err); + if (err) + return err; } else if (IS_ENABLED(CONFIG_NF_CONNTRACK_LABELS) && labels_nonzero(&info->labels.mask)) { err = ovs_ct_set_labels(ct, key, &info->labels.value, diff --git a/net/packet/af_packet.c b/net/packet/af_packet.c index 50cae32ae269..7c83e01526ed 100644 --- a/net/packet/af_packet.c +++ b/net/packet/af_packet.c @@ -617,7 +617,7 @@ static int prb_calc_retire_blk_tmo(struct packet_sock *po, return DEFAULT_PRB_RETIRE_TOV; div = ecmd.base.speed / 1000; - mbits = (blk_size_in_bytes * 8) / (1024 * 1024); + mbits = (u64)blk_size_in_bytes * 8 / (1024 * 1024); if (div) mbits /= div; @@ -2530,26 +2530,6 @@ static int tpacket_rcv(struct sk_buff *skb, struct net_device *dev, goto drop_n_restore; } -static void tpacket_destruct_skb(struct sk_buff *skb) -{ - struct packet_sock *po = pkt_sk(skb->sk); - - if (likely(po->tx_ring.pg_vec)) { - void *ph; - __u32 ts; - - ph = skb_zcopy_get_nouarg(skb); - - ts = __packet_set_timestamp(po, ph, skb); - __packet_set_status(po, ph, TP_STATUS_AVAILABLE | ts); - - packet_dec_pending(&po->tx_ring); - complete(&po->skb_completion); - } - - sock_wfree(skb); -} - static int __packet_snd_vnet_parse(struct virtio_net_hdr *vnet_hdr, size_t len) { if ((vnet_hdr->flags & VIRTIO_NET_HDR_F_NEEDS_CSUM) && @@ -2589,27 +2569,56 @@ static int packet_snd_vnet_parse(struct msghdr *msg, size_t *len, return 0; } +struct tpacket_uarg { + struct ubuf_info ubuf; + struct packet_sock *po; + void *ph; +}; + +static void tpacket_ubuf_complete(struct sk_buff *skb, struct ubuf_info *uarg, + bool success) +{ + struct tpacket_uarg *tu = container_of(uarg, struct tpacket_uarg, ubuf); + struct packet_sock *po = tu->po; + void *ph = tu->ph; + __u32 ts; + + DEBUG_NET_WARN_ON_ONCE(!skb); + + if (!refcount_dec_and_test(&uarg->refcnt)) + return; + + ts = __packet_set_timestamp(po, ph, skb); + __packet_set_status(po, ph, TP_STATUS_AVAILABLE | ts); + + packet_dec_pending(&po->tx_ring); + complete(&po->skb_completion); + + kfree(tu); + sk_free(&po->sk); +} + +static const struct ubuf_info_ops tpacket_ubuf_ops = { + .complete = tpacket_ubuf_complete, +}; + static int tpacket_fill_skb(struct packet_sock *po, struct sk_buff *skb, - void *frame, struct net_device *dev, void *data, int tp_len, + struct net_device *dev, void *data, int tp_len, __be16 proto, unsigned char *addr, int hlen, int copylen, int hard_header_len, const struct sockcm_cookie *sockc) { - union tpacket_uhdr ph; int to_write, offset, len, nr_frags, len_max; struct socket *sock = po->sk.sk_socket; struct page *page; int err; - ph.raw = frame; - skb->protocol = proto; skb->dev = dev; skb->priority = sockc->priority; skb->mark = sockc->mark; skb_set_delivery_type_by_clockid(skb, sockc->transmit_time, po->sk.sk_clockid); skb_setup_tx_timestamp(skb, sockc); - skb_zcopy_set_nouarg(skb, ph.raw); skb_reserve(skb, hlen); skb_reset_network_header(skb); @@ -2749,6 +2758,7 @@ static int tpacket_snd(struct packet_sock *po, struct msghdr *msg) struct virtio_net_hdr vnet_hdr; bool has_vnet_hdr = false; struct sockcm_cookie sockc; + struct tpacket_uarg *uarg; __be16 proto; int err, reserve = 0; void *ph; @@ -2876,7 +2886,7 @@ static int tpacket_snd(struct packet_sock *po, struct msghdr *msg) err = len_sum; goto out_status; } - tp_len = tpacket_fill_skb(po, skb, ph, dev, data, tp_len, proto, + tp_len = tpacket_fill_skb(po, skb, dev, data, tp_len, proto, addr, hlen, copylen, hard_header_len, &sockc); if (likely(tp_len >= 0) && @@ -2908,7 +2918,24 @@ static int tpacket_snd(struct packet_sock *po, struct msghdr *msg) virtio_net_hdr_set_proto(skb, &vnet_hdr); } - skb->destructor = tpacket_destruct_skb; + uarg = kmalloc(sizeof(*uarg), GFP_KERNEL); + if (unlikely(!uarg)) { + if (likely(len_sum > 0)) + err = len_sum; + else + err = -ENOMEM; + goto out_status; + } + uarg->po = po; + uarg->ph = ph; + uarg->ubuf.ops = &tpacket_ubuf_ops; + uarg->ubuf.flags = SKBFL_ZEROCOPY_FRAG; + refcount_set(&uarg->ubuf.refcnt, 1); + + /* Hold a sk_wmem_alloc reference until completion */ + refcount_inc(&po->sk.sk_wmem_alloc); + skb_zcopy_init(skb, &uarg->ubuf); + __packet_set_status(po, ph, TP_STATUS_SENDING); packet_inc_pending(&po->tx_ring); @@ -4486,21 +4513,20 @@ static struct pgv *alloc_pg_vec(struct tpacket_req *req, int order, bool tx_ring vec->len = block_nr; pg_vec = vec->pg_vec; + if (tx_ring) { + vec->deferred = kzalloc_obj(*vec->deferred, + GFP_KERNEL | __GFP_NOWARN); + if (!vec->deferred) + goto out_free_pgvec; + vec->deferred->vec = vec; + INIT_DELAYED_WORK(&vec->deferred->work, + packet_free_pg_vec_work); + } + for (i = 0; i < block_nr; i++) { pg_vec[i].buffer = alloc_one_pg_vec_page(order); if (unlikely(!pg_vec[i].buffer)) goto out_free_pgvec; - - if (tx_ring && !vec->deferred && - is_vmalloc_addr(pg_vec[i].buffer)) { - vec->deferred = kzalloc_obj(*vec->deferred, - GFP_KERNEL | __GFP_NOWARN); - if (!vec->deferred) - goto out_free_pgvec; - vec->deferred->vec = vec; - INIT_DELAYED_WORK(&vec->deferred->work, - packet_free_pg_vec_work); - } } out: diff --git a/net/rds/connection.c b/net/rds/connection.c index b6c4beb50eaf..c752a8623cfc 100644 --- a/net/rds/connection.c +++ b/net/rds/connection.c @@ -276,6 +276,12 @@ static struct rds_connection *__rds_conn_create(struct net *net, conn->c_trans = trans; + /* The transport may just have been swapped for loopback; size the + * set of paths - which is also what rds_conn_destroy() tears down + * again - by the transport the connection actually uses. + */ + npaths = (trans->t_mp_capable ? RDS_MPATH_WORKERS : 1); + init_waitqueue_head(&conn->c_hs_waitq); for (i = 0; i < npaths; i++) { __rds_conn_path_init(conn, &conn->c_path[i], diff --git a/net/rds/ib_frmr.c b/net/rds/ib_frmr.c index bd861191157b..8397aa4a17ac 100644 --- a/net/rds/ib_frmr.c +++ b/net/rds/ib_frmr.c @@ -204,19 +204,16 @@ static int rds_ib_map_frmr(struct rds_ib_device *rds_ibdev, */ rds_ib_teardown_mr(ibmr); - ibmr->sg = sg; - ibmr->sg_len = sg_len; - ibmr->sg_dma_len = 0; frmr->sg_byte_len = 0; - WARN_ON(ibmr->sg_dma_len); - ibmr->sg_dma_len = ib_dma_map_sg(dev, ibmr->sg, ibmr->sg_len, + ibmr->sg_dma_len = ib_dma_map_sg(dev, sg, sg_len, DMA_BIDIRECTIONAL); if (unlikely(!ibmr->sg_dma_len)) { pr_warn("RDS/IB: %s failed!\n", __func__); return -EBUSY; } - frmr->sg_byte_len = 0; + ibmr->sg = sg; + ibmr->sg_len = sg_len; frmr->dma_npages = 0; len = 0; @@ -264,6 +261,8 @@ static int rds_ib_map_frmr(struct rds_ib_device *rds_ibdev, ib_dma_unmap_sg(rds_ibdev->dev, ibmr->sg, ibmr->sg_len, DMA_BIDIRECTIONAL); ibmr->sg_dma_len = 0; + ibmr->sg = NULL; + ibmr->sg_len = 0; return ret; } diff --git a/net/sched/act_api.c b/net/sched/act_api.c index 3f653721c45f..e45a63be397c 100644 --- a/net/sched/act_api.c +++ b/net/sched/act_api.c @@ -758,7 +758,7 @@ static int tcf_idr_delete_index(struct tcf_idrinfo *idrinfo, u32 index) mutex_lock(&idrinfo->lock); p = idr_find(&idrinfo->action_idr, index); - if (!p) { + if (IS_ERR_OR_NULL(p)) { mutex_unlock(&idrinfo->lock); return -ENOENT; } diff --git a/net/sched/act_ct.c b/net/sched/act_ct.c index 370085ab6ea4..1d62bb654149 100644 --- a/net/sched/act_ct.c +++ b/net/sched/act_ct.c @@ -432,11 +432,10 @@ static void tcf_ct_flow_table_add(struct tcf_ct_flow_table *ct_ft, if (test_and_set_bit(IPS_OFFLOAD_BIT, &ct->status)) return; + /* NULL if ct is dying (raced flush) or the atomic alloc failed. */ entry = flow_offload_alloc(ct); - if (!entry) { - WARN_ON_ONCE(1); + if (!entry) goto err_alloc; - } if (tcp) { ct->proto.tcp.seen[0].flags |= IP_CT_TCP_FLAG_BE_LIBERAL; @@ -980,11 +979,11 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, struct tcf_result *res) { struct net *net = dev_net(skb->dev); + bool cached, commit, clear, nat; enum ip_conntrack_info ctinfo; struct tcf_ct *c = to_ct(a); struct nf_conn *tmpl = NULL; struct nf_hook_state state; - bool cached, commit, clear; int nh_ofs, err, retval; struct tcf_ct_params *p; bool add_helper = false; @@ -999,6 +998,7 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, retval = p->action; commit = p->ct_action & TCA_CT_ACT_COMMIT; clear = p->ct_action & TCA_CT_ACT_CLEAR; + nat = p->ct_action & TCA_CT_ACT_NAT; tmpl = p->tmpl; tcf_lastuse_update(&c->tcf_tm); @@ -1047,6 +1047,19 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, * different zone. */ cached = tcf_ct_skb_nfct_cached(net, skb, p); + + /* If the ct entry is not confirmed and shared with some other skb, + * e.g., a cloned one, we can't just modify it with a commit or nat + * as we must not modify the extension set. Reset. + */ + if (cached && (commit || nat)) { + ct = nf_ct_get(skb, &ctinfo); + if (ct && !nf_ct_is_confirmed(ct) && nf_ct_shared(ct)) { + nf_reset_ct(skb); + cached = false; + } + } + if (!cached) { if (tcf_ct_flow_table_lookup(p, skb, family)) { skip_add = true; @@ -1084,25 +1097,31 @@ TC_INDIRECT_SCOPE int tcf_ct_act(struct sk_buff *skb, const struct tc_action *a, if (err) goto drop; add_helper = true; - if (p->ct_action & TCA_CT_ACT_NAT && !nfct_seqadj(ct)) { + if (nat && !nfct_seqadj(ct)) { if (!nfct_seqadj_ext_add(ct)) goto drop; } } - if (nf_ct_is_confirmed(ct) ? ((!cached && !skip_add) || add_helper) : commit) { - err = nf_ct_helper(skb, ct, ctinfo, family); - if (err != NF_ACCEPT) - goto nf_error; - } - if (commit) { tcf_ct_act_set_mark(ct, p->mark, p->mark_mask); tcf_ct_act_set_labels(ct, p->labels, p->labels_mask); if (!nf_ct_is_confirmed(ct)) nf_conn_act_ct_ext_add(skb, ct, ctinfo); + } + /* Run helpers for the connection if nf_conntrack_in() was executed + * or if we're about to commit. This has to be done after all the + * extensions are already added. + */ + if (nf_ct_is_confirmed(ct) ? ((!cached && !skip_add) || add_helper) : commit) { + err = nf_ct_helper(skb, ct, ctinfo, family); + if (err != NF_ACCEPT) + goto nf_error; + } + + if (commit) { /* This will take care of sending queued events * even if the connection is already confirmed. */ diff --git a/net/sched/act_gate.c b/net/sched/act_gate.c index fdbfcaa3e2ab..a3c96519936b 100644 --- a/net/sched/act_gate.c +++ b/net/sched/act_gate.c @@ -681,7 +681,35 @@ static void tcf_gate_stats_update(struct tc_action *a, u64 bytes, u64 packets, static size_t tcf_gate_get_fill_size(const struct tc_action *act) { - return nla_total_size(sizeof(struct tc_gate)); + struct tcf_gate *gact = to_gate(act); + const struct tcf_gate_params *p; + struct tcfg_gate_entry *entry; + size_t size = nla_total_size(sizeof(struct tc_gate)) /* TCA_GATE_PARMS */ + + 3 * nla_total_size_64bit(sizeof(u64)) /* TCA_GATE_BASE_TIME + * TCA_GATE_CYCLE_TIME + * TCA_GATE_CYCLE_TIME_EXT + */ + + nla_total_size(sizeof(s32)) /* TCA_GATE_CLOCKID */ + + nla_total_size(sizeof(u32)) /* TCA_GATE_FLAGS */ + + nla_total_size(sizeof(s32)) /* TCA_GATE_PRIORITY */ + + nla_total_size(0); /* TCA_GATE_ENTRY_LIST */ + /* TCA_GATE_TM is budgeted by tcf_action_shared_attrs_size() */ + + rcu_read_lock(); + p = rcu_dereference(gact->param); + if (p) { + list_for_each_entry_rcu(entry, &p->entries, list) + /* TCA_GATE_ONE_ENTRY nest and its attributes */ + size += nla_total_size(0) + + nla_total_size(sizeof(u32)) /* TCA_GATE_ENTRY_INDEX */ + + nla_total_size(0) /* TCA_GATE_ENTRY_GATE */ + + nla_total_size(sizeof(u32)) /* TCA_GATE_ENTRY_INTERVAL */ + + nla_total_size(sizeof(s32)) /* TCA_GATE_ENTRY_MAX_OCTETS */ + + nla_total_size(sizeof(s32)); /* TCA_GATE_ENTRY_IPV */ + } + rcu_read_unlock(); + + return size; } static void tcf_gate_entry_destructor(void *priv) diff --git a/net/sched/act_ife.c b/net/sched/act_ife.c index 9cea71fc1db3..2afd68983ece 100644 --- a/net/sched/act_ife.c +++ b/net/sched/act_ife.c @@ -737,6 +737,7 @@ static int tcf_ife_decode(struct sk_buff *skb, const struct tc_action *a, u8 *curr_data; u16 mtype; u16 dlen; + int ret; curr_data = ife_tlv_meta_decode(tlv_data, ifehdr_end, &mtype, &dlen, NULL); @@ -745,13 +746,19 @@ static int tcf_ife_decode(struct sk_buff *skb, const struct tc_action *a, return TC_ACT_SHOT; } - if (find_decode_metaid(skb, p, mtype, dlen, curr_data)) { - /* abuse overlimits to count when we receive metadata - * but dont have an ops for it + ret = find_decode_metaid(skb, p, mtype, dlen, curr_data); + if (ret < 0) { + /* abuse overlimits to count metadata we cannot + * decode: no ops for it, or the decoder rejected it */ - pr_info_ratelimited("Unknown metaid %d dlen %d\n", - mtype, dlen); qstats_cpu_overlimit_inc(ife->common.cpu_qstats); + + if (ret == -ENOENT) + pr_info_ratelimited("Unknown metaid %d dlen %d\n", + mtype, dlen); + else + pr_info_ratelimited("Failed to decode metaid %d dlen %d err %d\n", + mtype, dlen, ret); } } diff --git a/net/sched/act_meta_mark.c b/net/sched/act_meta_mark.c index ea0573cb8b2d..e2f61b22bf0f 100644 --- a/net/sched/act_meta_mark.c +++ b/net/sched/act_meta_mark.c @@ -10,6 +10,7 @@ #include #include #include +#include #include #include #include @@ -28,9 +29,10 @@ static int skbmark_encode(struct sk_buff *skb, void *skbdata, static int skbmark_decode(struct sk_buff *skb, void *data, u16 len) { - u32 ifemark = *(u32 *)data; + if (len != sizeof(u32)) + return -EINVAL; - skb->mark = ntohl(ifemark); + skb->mark = get_unaligned_be32(data); return 0; } diff --git a/net/sched/act_meta_skbprio.c b/net/sched/act_meta_skbprio.c index 2df3133ce5ad..5cdb57931eab 100644 --- a/net/sched/act_meta_skbprio.c +++ b/net/sched/act_meta_skbprio.c @@ -10,6 +10,7 @@ #include #include #include +#include #include #include #include @@ -33,9 +34,10 @@ static int skbprio_encode(struct sk_buff *skb, void *skbdata, static int skbprio_decode(struct sk_buff *skb, void *data, u16 len) { - u32 ifeprio = *(u32 *)data; + if (len != sizeof(u32)) + return -EINVAL; - skb->priority = ntohl(ifeprio); + skb->priority = get_unaligned_be32(data); return 0; } diff --git a/net/sched/act_meta_skbtcindex.c b/net/sched/act_meta_skbtcindex.c index 44547caead46..8803710c0905 100644 --- a/net/sched/act_meta_skbtcindex.c +++ b/net/sched/act_meta_skbtcindex.c @@ -10,6 +10,7 @@ #include #include #include +#include #include #include #include @@ -28,9 +29,10 @@ static int skbtcindex_encode(struct sk_buff *skb, void *skbdata, static int skbtcindex_decode(struct sk_buff *skb, void *data, u16 len) { - u16 ifetc_index = *(u16 *)data; + if (len != sizeof(u16)) + return -EINVAL; - skb->tc_index = ntohs(ifetc_index); + skb->tc_index = get_unaligned_be16(data); return 0; } diff --git a/net/sched/cls_u32.c b/net/sched/cls_u32.c index a3e65c8cf29e..76ce2d124079 100644 --- a/net/sched/cls_u32.c +++ b/net/sched/cls_u32.c @@ -1003,8 +1003,16 @@ static int u32_change(struct net *net, struct sk_buff *in_skb, return -ENOMEM; } } else { - err = idr_alloc_u32(&tp_c->handle_idr, ht, &handle, - handle, GFP_KERNEL); + /* The IDR is keyed on the mapped id, and that is + * what the destroy paths remove. Ask for it here, + * so a manual handle colliding with the + * auto-allocated id space is rejected (-ENOSPC) + * instead of aliasing a future auto id. + */ + u32 id = handle2id(handle); + + err = idr_alloc_u32(&tp_c->handle_idr, ht, &id, id, + GFP_KERNEL); if (err) { kfree(ht); return err; diff --git a/net/sched/em_text.c b/net/sched/em_text.c index 343f1aebeec2..4132f8c3c5fc 100644 --- a/net/sched/em_text.c +++ b/net/sched/em_text.c @@ -113,7 +113,7 @@ static void em_text_destroy(struct tcf_ematch *m) static int em_text_dump(struct sk_buff *skb, struct tcf_ematch *m) { struct text_match *tm = EM_TEXT_PRIV(m); - struct tcf_em_text conf; + struct tcf_em_text conf = {}; strscpy(conf.algo, tm->config->ops->name); conf.from_offset = tm->from_offset; diff --git a/net/sched/sch_hfsc.c b/net/sched/sch_hfsc.c index e87f5021a199..284490fd6ca9 100644 --- a/net/sched/sch_hfsc.c +++ b/net/sched/sch_hfsc.c @@ -386,6 +386,15 @@ cftree_update(struct hfsc_class *cl) #define SM_MASK ((1ULL << SM_SHIFT) - 1) #define ISM_MASK ((1ULL << ISM_SHIFT) - 1) +/* + * Cap on the non-descending hops a classify walk may take before its + * filter chain is treated as misconfigured. A flowid binding that was + * legal at bind time can become lateral once hfsc_adjust_levels() + * raises a class level; a few such hops are legitimate, an unbounded + * run means the chain cycles. + */ +#define HFSC_CLASSIFY_MAX_DRIFT 8 + static inline u64 seg_x2y(u64 x, u64 sm) { @@ -1133,6 +1142,7 @@ hfsc_classify(struct sk_buff *skb, struct Qdisc *sch, int *qerr) struct hfsc_class *head, *cl; struct tcf_result res; struct tcf_proto *tcf; + unsigned int drift; int result; if (TC_H_MAJ(skb->priority ^ sch->handle) == 0 && @@ -1142,6 +1152,7 @@ hfsc_classify(struct sk_buff *skb, struct Qdisc *sch, int *qerr) *qerr = NET_XMIT_SUCCESS | __NET_XMIT_BYPASS; head = &q->root; + drift = HFSC_CLASSIFY_MAX_DRIFT; tcf = rcu_dereference_bh(q->root.filter_list); while (tcf && (result = tcf_classify_qdisc(skb, tcf, &res, false)) >= 0) { #ifdef CONFIG_NET_CLS_ACT @@ -1167,6 +1178,17 @@ hfsc_classify(struct sk_buff *skb, struct Qdisc *sch, int *qerr) if (cl->level == 0) return cl; /* hit leaf class */ + /* + * flowid binds skip the level check above (res.class is set + * at bind time and levels drift after), so a walk can follow + * lateral hops without descending; a bounded number of them + * is legal, more means the chain cycles. + */ + if (cl->level >= head->level && drift-- == 0) { + pr_warn_ratelimited("hfsc: classify hop budget exhausted, dropping packet\n"); + return NULL; + } + /* apply inner filter chain */ tcf = rcu_dereference_bh(cl->filter_list); head = cl; diff --git a/net/sched/sch_teql.c b/net/sched/sch_teql.c index 9e52afc2d980..409ce50cc0db 100644 --- a/net/sched/sch_teql.c +++ b/net/sched/sch_teql.c @@ -265,14 +265,11 @@ __teql_resolve(struct sk_buff *skb, struct sk_buff *skb_res, } if (neigh_event_send(n, skb_res) == 0) { - int err; char haddr[MAX_ADDR_LEN]; neigh_ha_snapshot(haddr, n, dev); - err = dev_hard_header(skb, dev, ntohs(skb_protocol(skb, false)), - haddr, NULL, skb->len); - - if (err < 0) + if (dev_hard_header(skb, dev, ntohs(skb_protocol(skb, false)), + haddr, NULL, skb->len) < 0) err = -EINVAL; } else { err = (skb_res == NULL) ? -EAGAIN : 1; diff --git a/net/sctp/associola.c b/net/sctp/associola.c index 5be0bed2685e..1220ff909fbe 100644 --- a/net/sctp/associola.c +++ b/net/sctp/associola.c @@ -1285,18 +1285,19 @@ void sctp_assoc_update_retran_path(struct sctp_association *asoc) /* Manually skip the head element. */ if (&trans->transports == &asoc->peer.transport_addr_list) continue; - if (trans->state == SCTP_UNCONFIRMED) - continue; - trans_next = sctp_trans_elect_best(trans, trans_next); - /* Active is good enough for immediate return. */ - if (trans_next->state == SCTP_ACTIVE) - break; + if (trans->state != SCTP_UNCONFIRMED) { + trans_next = sctp_trans_elect_best(trans, trans_next); + /* Active is good enough for immediate return. */ + if (trans_next->state == SCTP_ACTIVE) + break; + } /* We've reached the end, time to update path. */ if (trans == asoc->peer.retran_path) break; } - asoc->peer.retran_path = trans_next; + if (trans_next) + asoc->peer.retran_path = trans_next; pr_debug("%s: association:%p updated new path to addr:%pISpc\n", __func__, asoc, &asoc->peer.retran_path->ipaddr.sa); diff --git a/net/sctp/input.c b/net/sctp/input.c index 864741fae418..9494cfa51106 100644 --- a/net/sctp/input.c +++ b/net/sctp/input.c @@ -436,9 +436,10 @@ void sctp_icmp_proto_unreachable(struct sock *sk, if (timer_pending(&t->proto_unreach_timer)) return; else { - if (!mod_timer(&t->proto_unreach_timer, - jiffies + (HZ/20))) - sctp_transport_hold(t); + sctp_transport_hold(t); + if (mod_timer(&t->proto_unreach_timer, + jiffies + (HZ / 20))) + sctp_transport_put(t); } } else { struct net *net = sock_net(sk); diff --git a/net/sctp/sm_sideeffect.c b/net/sctp/sm_sideeffect.c index 0d99b7e8c082..35f540fb15fc 100644 --- a/net/sctp/sm_sideeffect.c +++ b/net/sctp/sm_sideeffect.c @@ -244,8 +244,9 @@ void sctp_generate_t3_rtx_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->T3_rtx_timer, jiffies + (HZ/20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->T3_rtx_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -280,8 +281,9 @@ static void sctp_generate_timeout_event(struct sctp_association *asoc, timeout_type); /* Try again later. */ - if (!mod_timer(&asoc->timers[timeout_type], jiffies + (HZ/20))) - sctp_association_hold(asoc); + sctp_association_hold(asoc); + if (mod_timer(&asoc->timers[timeout_type], jiffies + (HZ / 20))) + sctp_association_put(asoc); goto out_unlock; } @@ -378,8 +380,9 @@ void sctp_generate_heartbeat_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->hb_timer, jiffies + (HZ/20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->hb_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -388,8 +391,9 @@ void sctp_generate_heartbeat_event(struct timer_list *t) timeout = sctp_transport_timeout(transport); if (elapsed < timeout) { elapsed = timeout - elapsed; - if (!mod_timer(&transport->hb_timer, jiffies + elapsed)) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->hb_timer, jiffies + elapsed)) + sctp_transport_put(transport); goto out_unlock; } @@ -422,9 +426,10 @@ void sctp_generate_proto_unreach_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->proto_unreach_timer, - jiffies + (HZ/20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->proto_unreach_timer, + jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -458,8 +463,9 @@ void sctp_generate_reconf_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->reconf_timer, jiffies + (HZ / 20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->reconf_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } @@ -495,8 +501,9 @@ void sctp_generate_probe_event(struct timer_list *t) pr_debug("%s: sock is busy\n", __func__); /* Try again later. */ - if (!mod_timer(&transport->probe_timer, jiffies + (HZ / 20))) - sctp_transport_hold(transport); + sctp_transport_hold(transport); + if (mod_timer(&transport->probe_timer, jiffies + (HZ / 20))) + sctp_transport_put(transport); goto out_unlock; } diff --git a/net/sctp/sm_statefuns.c b/net/sctp/sm_statefuns.c index 3a8e16b29660..86f26b8b3993 100644 --- a/net/sctp/sm_statefuns.c +++ b/net/sctp/sm_statefuns.c @@ -2654,6 +2654,8 @@ static enum sctp_disposition sctp_sf_do_5_2_6_stale( sctp_add_cmd_sf(commands, SCTP_CMD_REPLY, SCTP_CHUNK(reply)); + sctp_add_cmd_sf(commands, SCTP_CMD_DISCARD_PACKET, SCTP_NULL()); + return SCTP_DISPOSITION_CONSUME; nomem: diff --git a/net/smc/smc_core.c b/net/smc/smc_core.c index 04aedd957543..9974149659c2 100644 --- a/net/smc/smc_core.c +++ b/net/smc/smc_core.c @@ -1849,6 +1849,7 @@ void smcr_port_err(struct smc_ib_device *smcibdev, u8 ibport) struct smc_link_group *lgr, *n; int i; + spin_lock_bh(&smc_lgr_list.lock); list_for_each_entry_safe(lgr, n, &smc_lgr_list.list, list) { if (strncmp(smcibdev->pnetid[ibport - 1], lgr->pnet_id, SMC_MAX_PNETID_LEN)) @@ -1863,6 +1864,7 @@ void smcr_port_err(struct smc_ib_device *smcibdev, u8 ibport) smcr_link_down_cond_sched(lnk); } } + spin_unlock_bh(&smc_lgr_list.lock); } static void smc_link_down_work(struct work_struct *work) diff --git a/net/smc/smc_ib.c b/net/smc/smc_ib.c index 9bb495707445..daaa8a72da90 100644 --- a/net/smc/smc_ib.c +++ b/net/smc/smc_ib.c @@ -333,6 +333,7 @@ static bool smc_ib_check_link_gid(u8 gid[SMC_GID_SIZE], bool smcrv2, static void smc_ib_gid_check(struct smc_ib_device *smcibdev, u8 ibport) { struct smc_link_group *lgr; + bool stale_gid = false; int i; spin_lock_bh(&smc_lgr_list.lock); @@ -348,11 +349,16 @@ static void smc_ib_gid_check(struct smc_ib_device *smcibdev, u8 ibport) continue; if (!smc_ib_check_link_gid(lgr->lnk[i].gid, lgr->smc_version == SMC_V2, - smcibdev, ibport)) - smcr_port_err(smcibdev, ibport); + smcibdev, ibport)) { + stale_gid = true; + goto out; + } } } +out: spin_unlock_bh(&smc_lgr_list.lock); + if (stale_gid) + smcr_port_err(smcibdev, ibport); } static int smc_ib_remember_port_attr(struct smc_ib_device *smcibdev, u8 ibport) diff --git a/net/tipc/group.c b/net/tipc/group.c index 14e6732624e2..74f6d3dac078 100644 --- a/net/tipc/group.c +++ b/net/tipc/group.c @@ -797,10 +797,10 @@ void tipc_group_proto_rcv(struct tipc_group *grp, bool *usr_wakeup, tipc_group_open(m, usr_wakeup); return; case GRP_ACK_MSG: - if (!m) + if (!m || !grp->bc_ackers) return; acked = msg_grp_bc_acked(hdr); - if (less_eq(acked, m->bc_acked)) + if (acked != grp->bc_snd_nxt || m->bc_acked == acked) return; m->bc_acked = acked; if (--grp->bc_ackers) diff --git a/net/tipc/monitor.c b/net/tipc/monitor.c index a94b9b36a700..1a438e312d66 100644 --- a/net/tipc/monitor.c +++ b/net/tipc/monitor.c @@ -632,9 +632,10 @@ static void mon_timeout(struct timer_list *t) { struct tipc_monitor *mon = timer_container_of(mon, t, timer); struct tipc_peer *self; - int best_member_cnt = dom_size(mon->peer_cnt) - 1; + int best_member_cnt; write_lock_bh(&mon->lock); + best_member_cnt = dom_size(mon->peer_cnt) - 1; self = mon->self; if (self && (best_member_cnt != self->applied)) { mon_update_local_domain(mon); diff --git a/net/vmw_vsock/af_vsock.c b/net/vmw_vsock/af_vsock.c index 78e60ad04363..4a1dfde8414f 100644 --- a/net/vmw_vsock/af_vsock.c +++ b/net/vmw_vsock/af_vsock.c @@ -2873,6 +2873,9 @@ static int vsock_net_child_mode_string(const struct ctl_table *table, int write, net = container_of(table->data, struct net, vsock.child_ns_mode); + if (!*lenp) + return 0; + ret = __vsock_net_mode_string(table, write, buffer, lenp, ppos, vsock_net_child_mode(net), &new_mode); if (ret) diff --git a/net/xdp/xskmap.c b/net/xdp/xskmap.c index 3bff346308d0..bf00d6463c19 100644 --- a/net/xdp/xskmap.c +++ b/net/xdp/xskmap.c @@ -124,7 +124,7 @@ static int xsk_map_gen_lookup(struct bpf_map *map, struct bpf_insn *insn_buf) struct bpf_insn *insn = insn_buf; *insn++ = BPF_LDX_MEM(BPF_W, ret, index, 0); - *insn++ = BPF_JMP_IMM(BPF_JGE, ret, map->max_entries, 5); + *insn++ = BPF_JMP32_IMM(BPF_JGE, ret, map->max_entries, 5); *insn++ = BPF_ALU64_IMM(BPF_LSH, ret, ilog2(sizeof(struct xsk_sock *))); *insn++ = BPF_ALU64_IMM(BPF_ADD, mp, offsetof(struct xsk_map, xsk_map)); *insn++ = BPF_ALU64_REG(BPF_ADD, ret, mp); diff --git a/security/ipe/eval.c b/security/ipe/eval.c index 21439c5be336..4a7be96c84f0 100644 --- a/security/ipe/eval.c +++ b/security/ipe/eval.c @@ -134,10 +134,14 @@ static bool evaluate_boot_verified(const struct ipe_eval_ctx *const ctx) static bool evaluate_dmv_roothash(const struct ipe_eval_ctx *const ctx, struct ipe_prop *p) { - return !!ctx->ipe_bdev && - !!ctx->ipe_bdev->root_hash && - ipe_digest_eval(p->value, - ctx->ipe_bdev->root_hash); + const struct digest_info *root_hash; + + if (!ctx->ipe_bdev) + return false; + + root_hash = rcu_dereference(ctx->ipe_bdev->root_hash); + + return root_hash && ipe_digest_eval(p->value, root_hash); } #else static bool evaluate_dmv_roothash(const struct ipe_eval_ctx *const ctx, diff --git a/security/ipe/eval.h b/security/ipe/eval.h index fef65a36468c..b26d1147d070 100644 --- a/security/ipe/eval.h +++ b/security/ipe/eval.h @@ -27,7 +27,7 @@ struct ipe_bdev { #ifdef CONFIG_IPE_PROP_DM_VERITY_SIGNATURE bool dm_verity_signed; #endif /* CONFIG_IPE_PROP_DM_VERITY_SIGNATURE */ - struct digest_info *root_hash; + struct digest_info __rcu *root_hash; }; #endif /* CONFIG_IPE_PROP_DM_VERITY */ diff --git a/security/ipe/fs.c b/security/ipe/fs.c index 076c111c85c8..847a76afb93d 100644 --- a/security/ipe/fs.c +++ b/security/ipe/fs.c @@ -159,18 +159,16 @@ static ssize_t new_policy(struct file *f, const char __user *data, } rc = ipe_new_policyfs_node(p); - if (rc) - goto out; out: kfree(copy); if (rc < 0) { ipe_free_policy(p); ipe_audit_policy_load(ERR_PTR(rc)); - } else { - ipe_audit_policy_load(p); + return rc; } - return (rc < 0) ? rc : len; + + return len; } static const struct file_operations np_fops = { diff --git a/security/ipe/hooks.c b/security/ipe/hooks.c index 0ae54a880405..d878b52ceec8 100644 --- a/security/ipe/hooks.c +++ b/security/ipe/hooks.c @@ -9,6 +9,7 @@ #include #include #include +#include #include "ipe.h" #include "hooks.h" @@ -232,7 +233,20 @@ void ipe_bdev_free_security(struct block_device *bdev) { struct ipe_bdev *blob = ipe_bdev(bdev); - ipe_digest_free(blob->root_hash); + ipe_digest_free(rcu_access_pointer(blob->root_hash)); +} + +static void ipe_set_dmverity_roothash(struct ipe_bdev *blob, + struct digest_info *info) +{ + struct digest_info *old; + + /* Protected by device-mapper's md->suspend_lock */ + old = rcu_replace_pointer(blob->root_hash, info, true); + if (old) { + synchronize_rcu(); + ipe_digest_free(old); + } } #ifdef CONFIG_IPE_PROP_DM_VERITY_SIGNATURE @@ -280,8 +294,7 @@ int ipe_bdev_setintegrity(struct block_device *bdev, enum lsm_integrity_type typ return -EINVAL; if (!value) { - ipe_digest_free(blob->root_hash); - blob->root_hash = NULL; + ipe_set_dmverity_roothash(blob, NULL); return 0; } @@ -301,8 +314,7 @@ int ipe_bdev_setintegrity(struct block_device *bdev, enum lsm_integrity_type typ info->digest_len = digest->digest_len; - ipe_digest_free(blob->root_hash); - blob->root_hash = info; + ipe_set_dmverity_roothash(blob, info); return 0; err: diff --git a/security/ipe/policy_fs.c b/security/ipe/policy_fs.c index 9d92d8a14b13..a7aeb57483c6 100644 --- a/security/ipe/policy_fs.c +++ b/security/ipe/policy_fs.c @@ -481,6 +481,9 @@ int ipe_new_policyfs_node(struct ipe_policy *p) inode_lock(root); p->policyfs = policyfs; root->i_private = p; + /* Only audit signed policies from userspace */ + if (p->pkcs7) + ipe_audit_policy_load(p); inode_unlock(root); return 0; diff --git a/security/landlock/access.h b/security/landlock/access.h index d926078bf0a5..a0fa1009ccdb 100644 --- a/security/landlock/access.h +++ b/security/landlock/access.h @@ -61,6 +61,10 @@ union access_masks_all { static_assert(sizeof(typeof_member(union access_masks_all, masks)) == sizeof(typeof_member(union access_masks_all, all))); +#define _LANDLOCK_LAYER_MASK_PADDING \ + (BITS_PER_TYPE(access_mask_t) - LANDLOCK_NUM_ACCESS_MAX - \ + IS_ENABLED(CONFIG_AUDIT)) + /** * struct layer_mask - The access rights and rule flags for a layer. * @@ -81,6 +85,10 @@ struct layer_mask { */ access_mask_t quiet : 1; #endif /* CONFIG_AUDIT */ + /** + * @__pad: Padding for the compiler's bitfield initialization. + */ + access_mask_t __pad : _LANDLOCK_LAYER_MASK_PADDING; } __packed __aligned(sizeof(access_mask_t)); /* diff --git a/sound/hda/codecs/realtek/alc269.c b/sound/hda/codecs/realtek/alc269.c index c31fed171c41..1e0612c39c89 100644 --- a/sound/hda/codecs/realtek/alc269.c +++ b/sound/hda/codecs/realtek/alc269.c @@ -4209,6 +4209,7 @@ enum { ALC245_FIXUP_CLEVO_NOISY_MIC, ALC269_FIXUP_VAIO_VJFH52_MIC_NO_PRESENCE, ALC233_FIXUP_MEDION_MTL_SPK, + ALC269_FIXUP_STARLABS_LIMIT_INT_MIC_BOOST, ALC233_FIXUP_STARLABS_STARFIGHTER, ALC294_FIXUP_BASS_SPEAKER_15, ALC283_FIXUP_DELL_HP_RESUME, @@ -6755,9 +6756,15 @@ static const struct hda_fixup alc269_fixups[] = { { } }, }, + [ALC269_FIXUP_STARLABS_LIMIT_INT_MIC_BOOST] = { + .type = HDA_FIXUP_FUNC, + .v.func = alc269_fixup_limit_int_mic_boost, + }, [ALC233_FIXUP_STARLABS_STARFIGHTER] = { .type = HDA_FIXUP_FUNC, .v.func = alc233_fixup_starlabs_starfighter, + .chained = true, + .chain_id = ALC269_FIXUP_STARLABS_LIMIT_INT_MIC_BOOST, }, [ALC294_FIXUP_BASS_SPEAKER_15] = { .type = HDA_FIXUP_FUNC, @@ -8085,6 +8092,7 @@ static const struct hda_quirk alc269_fixup_tbl[] = { SND_PCI_QUIRK(0x1f66, 0x0105, "Ayaneo Portable Game Player", ALC287_FIXUP_CS35L41_I2C_2), SND_PCI_QUIRK(0x2014, 0x800a, "Positivo ARN50", ALC269_FIXUP_LIMIT_INT_MIC_BOOST), SND_PCI_QUIRK(0x2039, 0x0001, "Inspur S14-G1", ALC295_FIXUP_CHROME_BOOK), + SND_PCI_QUIRK(0x2145, 0x0001, "Star Labs StarFighter", ALC233_FIXUP_STARLABS_STARFIGHTER), SND_PCI_QUIRK(0x2782, 0x0214, "VAIO VJFE-CL", ALC269_FIXUP_LIMIT_INT_MIC_BOOST), SND_PCI_QUIRK(0x2782, 0x0228, "Infinix ZERO BOOK 13", ALC269VB_FIXUP_INFINIX_ZERO_BOOK_13), SND_PCI_QUIRK(0x2782, 0x0232, "CHUWI CoreBook XPro", ALC269VB_FIXUP_CHUWI_COREBOOK_XPRO), @@ -8168,6 +8176,7 @@ static const struct hda_quirk alc269_fixup_vendor_tbl[] = { SND_PCI_QUIRK_VENDOR(0x104d, "Sony VAIO", ALC269_FIXUP_SONY_VAIO), SND_PCI_QUIRK_VENDOR(0x17aa, "Lenovo XPAD", ALC269_FIXUP_LENOVO_XPAD_ACPI), SND_PCI_QUIRK_VENDOR(0x19e5, "Huawei Matebook", ALC255_FIXUP_MIC_MUTE_LED), + SND_PCI_QUIRK_VENDOR(0x2145, "Star Labs", ALC269_FIXUP_STARLABS_LIMIT_INT_MIC_BOOST), {} }; diff --git a/tools/arch/riscv/include/asm/csr.h b/tools/arch/riscv/include/asm/csr.h index 21d8cee04638..8df64314d613 100644 --- a/tools/arch/riscv/include/asm/csr.h +++ b/tools/arch/riscv/include/asm/csr.h @@ -163,12 +163,24 @@ #define HGATP_MODE_SHIFT HGATP32_MODE_SHIFT #endif -/* VSIP & HVIP relation */ +/* + * VSIP & HVIP relation + * + * The bit positions are same between VSIP and HVIP for interrupt + * numbers 13-63, where there's a shift for the SSI, STI and SEI. + */ #define VSIP_TO_HVIP_SHIFT (IRQ_VS_SOFT - IRQ_S_SOFT) -#define VSIP_VALID_MASK ((_AC(1, UL) << IRQ_S_SOFT) | \ +#define VSIP_BIAS_MASK ((_AC(1, UL) << IRQ_S_SOFT) | \ (_AC(1, UL) << IRQ_S_TIMER) | \ - (_AC(1, UL) << IRQ_S_EXT) | \ - (_AC(1, UL) << IRQ_PMU_OVF)) + (_AC(1, UL) << IRQ_S_EXT)) +#define VSIP_NO_BIAS_MASK (_AC(1, UL) << IRQ_PMU_OVF) +#define VSIP_VALID_MASK (VSIP_BIAS_MASK | VSIP_NO_BIAS_MASK) +#define vsip_to_hvip(_vsip) ((((_vsip) & VSIP_BIAS_MASK) << \ + VSIP_TO_HVIP_SHIFT) | \ + ((_vsip) & VSIP_NO_BIAS_MASK)) +#define hvip_to_vsip(_hvip) ((((_hvip) >> VSIP_TO_HVIP_SHIFT) & \ + VSIP_BIAS_MASK) | \ + ((_hvip) & VSIP_NO_BIAS_MASK)) /* AIA CSR bits */ #define TOPI_IID_SHIFT 16 diff --git a/tools/lib/bpf/libbpf.c b/tools/lib/bpf/libbpf.c index 1ab939dfb7f0..4f9ce4ff005b 100644 --- a/tools/lib/bpf/libbpf.c +++ b/tools/lib/bpf/libbpf.c @@ -6162,6 +6162,13 @@ bpf_object__relocate_core(struct bpf_object *obj, const char *targ_btf_path) return -EINVAL; insn = &prog->insns[insn_idx]; + if (is_ldimm64_insn(insn) && (size_t)insn_idx + 1 >= prog->insns_cnt) { + pr_warn("prog '%s': relo #%d: insn #%d (LDIMM64) is truncated\n", + prog->name, i, insn_idx); + err = -EINVAL; + goto out; + } + err = record_relo_core(prog, rec, insn_idx); if (err) { pr_warn("prog '%s': relo #%d: failed to record relocation: %s\n", diff --git a/tools/lib/bpf/relo_core.c b/tools/lib/bpf/relo_core.c index 6ae3f2a15ad0..4755de3f9995 100644 --- a/tools/lib/bpf/relo_core.c +++ b/tools/lib/bpf/relo_core.c @@ -980,23 +980,30 @@ static int bpf_core_calc_relo(const char *prog_name, } /* - * Turn instruction for which CO_RE relocation failed into invalid one with + * Turn instruction for which CO-RE relocation failed into invalid one with * distinct signature. */ -static void bpf_core_poison_insn(const char *prog_name, int relo_idx, - int insn_idx, struct bpf_insn *insn) +static int bpf_core_poison_insn(const char *prog_name, int relo_idx, + struct bpf_insn *insn, int insn_idx) { - pr_debug("prog '%s': relo #%d: substituting insn #%d w/ invalid insn\n", - prog_name, relo_idx, insn_idx); - insn->code = BPF_JMP | BPF_CALL; - insn->dst_reg = 0; - insn->src_reg = 0; - insn->off = 0; - /* if this instruction is reachable (not a dead code), - * verifier will complain with the following message: - * invalid func unknown#195896080 - */ - insn->imm = 195896080; /* => 0xbad2310 => "bad relo" */ + int insn_cnt = is_ldimm64_insn(insn) ? 2 : 1; + int i; + + for (i = 0; i < insn_cnt; i++) { + pr_debug("prog '%s': relo #%d: substituting insn #%d w/ invalid insn\n", + prog_name, relo_idx, insn_idx + i); + insn[i].code = BPF_JMP | BPF_CALL; + insn[i].dst_reg = 0; + insn[i].src_reg = 0; + insn[i].off = 0; + /* + * If this instruction is reachable (not dead code), the verifier + * will complain with "invalid func unknown#195896080". + */ + insn[i].imm = 195896080; /* => 0xbad2310 => "bad relo" */ + } + + return 0; } static int insn_bpf_size_to_bytes(struct bpf_insn *insn) @@ -1047,17 +1054,6 @@ int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn, class = BPF_CLASS(insn->code); - if (res->poison) { -poison: - /* poison second part of ldimm64 to avoid confusing error from - * verifier about "unknown opcode 00" - */ - if (is_ldimm64_insn(insn)) - bpf_core_poison_insn(prog_name, relo_idx, insn_idx + 1, insn + 1); - bpf_core_poison_insn(prog_name, relo_idx, insn_idx, insn); - return 0; - } - orig_val = res->orig_val; new_val = res->new_val; @@ -1065,7 +1061,9 @@ int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn, case BPF_ALU: case BPF_ALU64: if (BPF_SRC(insn->code) != BPF_K) - return -EINVAL; + goto bad_insn; + if (res->poison) + return bpf_core_poison_insn(prog_name, relo_idx, insn, insn_idx); if (res->validate && insn->imm != orig_val) { pr_warn("prog '%s': relo #%d: unexpected insn #%d (ALU/ALU64) value: got %u, exp %llu -> %llu\n", prog_name, relo_idx, @@ -1082,6 +1080,8 @@ int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn, case BPF_LDX: case BPF_ST: case BPF_STX: + if (res->poison) + return bpf_core_poison_insn(prog_name, relo_idx, insn, insn_idx); if (res->validate && insn->off != orig_val) { pr_warn("prog '%s': relo #%d: unexpected insn #%d (LDX/ST/STX) value: got %u, exp %llu -> %llu\n", prog_name, relo_idx, insn_idx, insn->off, (unsigned long long)orig_val, @@ -1097,7 +1097,7 @@ int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn, pr_warn("prog '%s': relo #%d: insn #%d (LDX/ST/STX) accesses field incorrectly. " "Make sure you are accessing pointers, unsigned integers, or fields of matching type and size.\n", prog_name, relo_idx, insn_idx); - goto poison; + return bpf_core_poison_insn(prog_name, relo_idx, insn, insn_idx); } orig_val = insn->off; @@ -1140,6 +1140,9 @@ int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn, return -EINVAL; } + if (res->poison) + return bpf_core_poison_insn(prog_name, relo_idx, insn, insn_idx); + imm = (__u32)insn[0].imm | ((__u64)insn[1].imm << 32); if (res->validate && imm != orig_val) { pr_warn("prog '%s': relo #%d: unexpected insn #%d (LDIMM64) value: got %llu, exp %llu -> %llu\n", @@ -1157,6 +1160,7 @@ int bpf_core_patch_insn(const char *prog_name, struct bpf_insn *insn, break; } default: +bad_insn: pr_warn("prog '%s': relo #%d: trying to relocate unrecognized insn #%d, code:0x%x, src:0x%x, dst:0x%x, off:0x%x, imm:0x%x\n", prog_name, relo_idx, insn_idx, insn->code, insn->src_reg, insn->dst_reg, insn->off, insn->imm); diff --git a/tools/testing/selftests/bpf/prog_tests/linked_list.c b/tools/testing/selftests/bpf/prog_tests/linked_list.c index 8defea0253ed..178d133d5923 100644 --- a/tools/testing/selftests/bpf/prog_tests/linked_list.c +++ b/tools/testing/selftests/bpf/prog_tests/linked_list.c @@ -713,7 +713,7 @@ static void test_btf(void) break; err = btf__load_into_kernel(btf); - ASSERT_EQ(err, -ELOOP, "check btf"); + ASSERT_EQ(err, 0, "check btf"); btf__free(btf); break; } @@ -772,7 +772,7 @@ static void test_btf(void) break; err = btf__load_into_kernel(btf); - ASSERT_EQ(err, -ELOOP, "check btf"); + ASSERT_EQ(err, 0, "check btf"); btf__free(btf); break; } diff --git a/tools/testing/selftests/cgroup/test_cpuset_prs.sh b/tools/testing/selftests/cgroup/test_cpuset_prs.sh index ebbc5b4def24..2dff9d5b5356 100755 --- a/tools/testing/selftests/cgroup/test_cpuset_prs.sh +++ b/tools/testing/selftests/cgroup/test_cpuset_prs.sh @@ -298,6 +298,8 @@ TEST_MATRIX=( " C0-4:X2-4 C1-4:X2-4:P2 C2-4:X4:P1 \ . . . X1 . 0 A1:0-1|A2:2-4|A3:2-4 \ A1:P0|A2:P2|A3:P-1 2-4" + " CX1-3:P1 CX1-3 CX1-3 . . . P1 . 0 A1:1-3|A2:1-3|A3:1-3 \ + A1:P1|A2:P0|A3:P-1" # Remote partition offline tests " C0-3 C1-3 C2-3 . X2-3 X2-3 X2-3:P2:O2=0 . 0 A1:0-1|A2:1|A3:3 A1:P0|A3:P2 2-3" diff --git a/tools/testing/selftests/cgroup/test_memcontrol.c b/tools/testing/selftests/cgroup/test_memcontrol.c index 3a84d068fbf3..0ed82347044e 100644 --- a/tools/testing/selftests/cgroup/test_memcontrol.c +++ b/tools/testing/selftests/cgroup/test_memcontrol.c @@ -30,7 +30,7 @@ static int page_size; int get_temp_fd(void) { - return open(".", O_TMPFILE | O_RDWR | O_EXCL); + return open(".", O_TMPFILE | O_RDWR | O_EXCL, 0600); } int alloc_pagecache(int fd, size_t size) diff --git a/tools/testing/selftests/ftrace/test.d/kprobe/kprobe_non_uniq_symbol.tc b/tools/testing/selftests/ftrace/test.d/kprobe/kprobe_non_uniq_symbol.tc index bc9514428dba..07b1177c1634 100644 --- a/tools/testing/selftests/ftrace/test.d/kprobe/kprobe_non_uniq_symbol.tc +++ b/tools/testing/selftests/ftrace/test.d/kprobe/kprobe_non_uniq_symbol.tc @@ -6,7 +6,7 @@ SYMBOL='name_show' # We skip this test on kernel where SYMBOL is unique or does not exist. -if [ "$(grep -c -E "[[:alnum:]]+ t ${SYMBOL}" /proc/kallsyms)" -le '1' ]; then +if [ "$(grep -c -E "[[:alnum:]]+ t ${SYMBOL}$" /proc/kallsyms)" -le '1' ]; then exit_unsupported fi diff --git a/tools/testing/selftests/kvm/steal_time.c b/tools/testing/selftests/kvm/steal_time.c index 76fcdd1fd3cb..c00cb17f9e40 100644 --- a/tools/testing/selftests/kvm/steal_time.c +++ b/tools/testing/selftests/kvm/steal_time.c @@ -27,6 +27,9 @@ static void *st_gva[NR_VCPUS]; static u64 guest_stolen_time[NR_VCPUS]; +static struct kvm_vm *vm_create_steal_time(u32 nr_vcpus, void *guest_code, + struct kvm_vcpu *vcpus[]); + #if defined(__x86_64__) /* steal_time must have 64-byte alignment */ @@ -211,17 +214,14 @@ static void check_steal_time_uapi(void) u64 st_ipa; int ret; - vm = vm_create_with_one_vcpu(&vcpu, NULL); - struct kvm_device_attr dev = { .group = KVM_ARM_VCPU_PVTIME_CTRL, .attr = KVM_ARM_VCPU_PVTIME_IPA, .addr = (u64)&st_ipa, }; + vm = vm_create_steal_time(1, NULL, &vcpu); vcpu_ioctl(vcpu, KVM_HAS_DEVICE_ATTR, &dev); - vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, ST_GPA_BASE, 1, 1, 0); - virt_map(vm, ST_GPA_BASE, ST_GPA_BASE, 1); st_ipa = (ulong)ST_GPA_BASE | 1; ret = __vcpu_ioctl(vcpu, KVM_SET_DEVICE_ATTR, &dev); @@ -504,6 +504,21 @@ static void run_vcpu(struct kvm_vcpu *vcpu) } } +static struct kvm_vm *vm_create_steal_time(u32 nr_vcpus, void *guest_code, + struct kvm_vcpu *vcpus[]) +{ + unsigned int gpages; + struct kvm_vm *vm; + + /* Create a VM and an identity mapped memslot for the steal time structure */ + vm = vm_create_with_vcpus(nr_vcpus, guest_code, vcpus); + gpages = vm_calc_num_guest_pages(VM_MODE_DEFAULT, STEAL_TIME_SIZE * nr_vcpus); + vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, ST_GPA_BASE, 1, gpages, 0); + virt_map(vm, ST_GPA_BASE, ST_GPA_BASE, gpages); + + return vm; +} + int main(int ac, char **av) { struct kvm_vcpu *vcpus[NR_VCPUS]; @@ -511,7 +526,6 @@ int main(int ac, char **av) pthread_attr_t attr; pthread_t thread; cpu_set_t cpuset; - unsigned int gpages; long stolen_time; long run_delay; bool verbose; @@ -526,11 +540,7 @@ int main(int ac, char **av) pthread_attr_setaffinity_np(&attr, sizeof(cpu_set_t), &cpuset); pthread_setaffinity_np(pthread_self(), sizeof(cpu_set_t), &cpuset); - /* Create a VM and an identity mapped memslot for the steal time structure */ - vm = vm_create_with_vcpus(NR_VCPUS, guest_code, vcpus); - gpages = vm_calc_num_guest_pages(VM_MODE_DEFAULT, STEAL_TIME_SIZE * NR_VCPUS); - vm_userspace_mem_region_add(vm, VM_MEM_SRC_ANONYMOUS, ST_GPA_BASE, 1, gpages, 0); - virt_map(vm, ST_GPA_BASE, ST_GPA_BASE, gpages); + vm = vm_create_steal_time(NR_VCPUS, guest_code, vcpus); ksft_print_header(); TEST_REQUIRE(is_steal_time_supported(vcpus[0])); diff --git a/tools/testing/selftests/nci/nci_dev.c b/tools/testing/selftests/nci/nci_dev.c index 312f84ee0444..07427fa42888 100644 --- a/tools/testing/selftests/nci/nci_dev.c +++ b/tools/testing/selftests/nci/nci_dev.c @@ -8,6 +8,7 @@ #include #include +#include #include #include #include @@ -87,6 +88,16 @@ struct msgtemplate { char buf[MAX_MSG_SIZE]; }; +static int join_thread_status(pthread_t thread) +{ + void *thread_ret = NULL; + + if (pthread_join(thread, &thread_ret)) + return -1; + + return (int)(intptr_t)thread_ret; +} + static int create_nl_socket(void) { int fd; @@ -182,7 +193,7 @@ static int get_family_id(int sd, __u32 pid, __u32 *event_group) } ans; struct nlattr *na; int resp_len; - __u16 id; + __u16 id = 0; int len; int rc; @@ -438,13 +449,13 @@ FIXTURE_SETUP(NCI) else rc = pthread_create(&thread_t, NULL, virtual_dev_open, (void *)&self->virtual_nci_fd); - ASSERT_GT(rc, -1); + ASSERT_EQ(rc, 0); rc = send_cmd_with_idx(self->sd, self->fid, self->pid, NFC_CMD_DEV_UP, self->dev_idex); EXPECT_EQ(rc, 0); - pthread_join(thread_t, (void **)&status); + status = join_thread_status(thread_t); ASSERT_EQ(status, 0); self->open_state = true; } @@ -509,12 +520,12 @@ FIXTURE_TEARDOWN(NCI) rc = pthread_create(&thread_t, NULL, virtual_deinit, (void *)&self->virtual_nci_fd); - ASSERT_GT(rc, -1); + ASSERT_EQ(rc, 0); rc = send_cmd_with_idx(self->sd, self->fid, self->pid, NFC_CMD_DEV_DOWN, self->dev_idex); EXPECT_EQ(rc, 0); - pthread_join(thread_t, (void **)&status); + status = join_thread_status(thread_t); ASSERT_EQ(status, 0); } @@ -585,12 +596,11 @@ int start_polling(int dev_idx, int proto, int virtual_fd, int sd, int fid, int p void *nla_start_poll_data[2] = {&dev_idx, &proto}; int nla_start_poll_len[2] = {4, 4}; pthread_t thread_t; - int status; int rc; rc = pthread_create(&thread_t, NULL, virtual_poll_start, (void *)&virtual_fd); - if (rc < 0) + if (rc) return rc; rc = send_cmd_mt_nla(sd, fid, pid, NFC_CMD_START_POLL, 2, nla_start_poll_type, @@ -598,19 +608,17 @@ int start_polling(int dev_idx, int proto, int virtual_fd, int sd, int fid, int p if (rc != 0) return rc; - pthread_join(thread_t, (void **)&status); - return status; + return join_thread_status(thread_t); } int stop_polling(int dev_idx, int virtual_fd, int sd, int fid, int pid) { pthread_t thread_t; - int status; int rc; rc = pthread_create(&thread_t, NULL, virtual_poll_stop, (void *)&virtual_fd); - if (rc < 0) + if (rc) return rc; rc = send_cmd_with_idx(sd, fid, pid, @@ -618,8 +626,7 @@ int stop_polling(int dev_idx, int virtual_fd, int sd, int fid, int pid) if (rc != 0) return rc; - pthread_join(thread_t, (void **)&status); - return status; + return join_thread_status(thread_t); } TEST_F(NCI, start_poll) @@ -830,10 +837,14 @@ int disconnect_tag(int nfc_sock, int virtual_fd) status = pthread_create(&thread_t, NULL, virtual_deactivate_proc, (void *)&virtual_fd); + if (status) + return status; close(nfc_sock); - pthread_join(thread_t, (void **)&status); - return status; + if (status) + return -1; + + return join_thread_status(thread_t); } TEST_F(NCI, t4t_tag_read) @@ -874,13 +885,13 @@ TEST_F(NCI, deinit) else rc = pthread_create(&thread_t, NULL, virtual_deinit, (void *)&self->virtual_nci_fd); - ASSERT_GT(rc, -1); + ASSERT_EQ(rc, 0); rc = send_cmd_with_idx(self->sd, self->fid, self->pid, NFC_CMD_DEV_DOWN, self->dev_idex); EXPECT_EQ(rc, 0); - pthread_join(thread_t, (void **)&status); + status = join_thread_status(thread_t); self->open_state = 0; ASSERT_EQ(status, 0); diff --git a/virt/kvm/kvm_main.c b/virt/kvm/kvm_main.c index 45e784462ec6..e389ac4744e6 100644 --- a/virt/kvm/kvm_main.c +++ b/virt/kvm/kvm_main.c @@ -1116,7 +1116,7 @@ static struct kvm *kvm_create_vm(unsigned long type, const char *fdname) rcuwait_init(&kvm->mn_memslots_update_rcuwait); xa_init(&kvm->vcpu_array); #ifdef CONFIG_KVM_GENERIC_MEMORY_ATTRIBUTES - xa_init(&kvm->mem_attr_array); + xa_init_flags(&kvm->mem_attr_array, XA_FLAGS_ACCOUNT); #endif INIT_LIST_HEAD(&kvm->gpc_list); @@ -2446,14 +2446,36 @@ bool kvm_range_has_memory_attributes(struct kvm *kvm, gfn_t start, gfn_t end, return (kvm_get_memory_attributes(kvm, start) & mask) == attrs; guard(rcu)(); - if (!attrs) - return !xas_find(&xas, end - 1); + /* + * Lookup the entry for each index instead of iterating over the xarray + * as KVM deletes/nullifies entries to represent "no attributes", and + * the xas index is effectively invalid when no entry is found. I.e. + * matching non-zero attributes for *every* entry effectively requires + * a manually lookup for each index. + * + * Skip pre-allocated, reserved entries, or restart the lookup if the + * xarray was concurrently modified, via xas_retry() ("retry" means the + * entry holds an internal xarray value, i.e. is either invalid or NULL + * from the caller's perspective). + * + * Use xas_next() when looking for non-zero attributes to optimize for + * the case where the start of the range (or the entire range) doesn't + * have any attributes, as xas_next() returns literally the next entry, + * whereas xas_next_entry() returns the next non-NULL entry (bounded by + * a maximum index). + */ for (index = start; index < end; index++) { do { - entry = xas_next(&xas); + entry = attrs ? xas_next(&xas) : + xas_next_entry(&xas, end - 1); } while (xas_retry(&xas, entry)); + if (!entry) + return !attrs; + + WARN_ON_ONCE(!xa_to_value(entry)); + if (xas.xa_index != index || (xa_to_value(entry) & mask) != attrs) return false;