linux

mirror of https://github.com/torvalds/linux.git synced 2024-11-27 22:51:35 +00:00

History

Dave Marchevsky d2dcc67df9 bpf: Migrate bpf_rbtree_add and bpf_list_push_{front,back} to possibly fail Consider this code snippet: struct node { long key; bpf_list_node l; bpf_rb_node r; bpf_refcount ref; } int some_bpf_prog(void ctx) { struct node n = bpf_obj_new(/.../), m; bpf_spin_lock(&glock); bpf_rbtree_add(&some_tree, &n->r, / ... /); m = bpf_refcount_acquire(n); bpf_rbtree_add(&other_tree, &m->r, / ... /); bpf_spin_unlock(&glock); / ... / } After bpf_refcount_acquire, n and m point to the same underlying memory, and that node's bpf_rb_node field is being used by the some_tree insert, so overwriting it as a result of the second insert is an error. In order to properly support refcounted nodes, the rbtree and list insert functions must be allowed to fail. This patch adds such support. The kfuncs bpf_rbtree_add, bpf_list_push_{front,back} are modified to return an int indicating success/failure, with 0 -> success, nonzero -> failure. bpf_obj_drop on failure ======================= Currently the only reason an insert can fail is the example above: the bpf_{list,rb}_node is already in use. When such a failure occurs, the insert kfuncs will bpf_obj_drop the input node. This allows the insert operations to logically fail without changing their verifier owning ref behavior, namely the unconditional release_reference of the input owning ref. With insert that always succeeds, ownership of the node is always passed to the collection, since the node always ends up in the collection. With a possibly-failed insert w/ bpf_obj_drop, ownership of the node is always passed either to the collection (success), or to bpf_obj_drop (failure). Regardless, it's correct to continue unconditionally releasing the input owning ref, as something is always taking ownership from the calling program on insert. Keeping owning ref behavior unchanged results in a nice default UX for insert functions that can fail. If the program's reaction to a failed insert is "fine, just get rid of this owning ref for me and let me go on with my business", then there's no reason to check for failure since that's default behavior. e.g.: long important_failures = 0; int some_bpf_prog(void ctx) { struct node n, m, o; / all bpf_obj_new'd / bpf_spin_lock(&glock); bpf_rbtree_add(&some_tree, &n->node, / ... /); bpf_rbtree_add(&some_tree, &m->node, / ... /); if (bpf_rbtree_add(&some_tree, &o->node, / ... /)) { important_failures++; } bpf_spin_unlock(&glock); } If we instead chose to pass ownership back to the program on failed insert - by returning NULL on success or an owning ref on failure - programs would always have to do something with the returned ref on failure. The most likely action is probably "I'll just get rid of this owning ref and go about my business", which ideally would look like: if (n = bpf_rbtree_add(&some_tree, &n->node, / ... /)) bpf_obj_drop(n); But bpf_obj_drop isn't allowed in a critical section and inserts must occur within one, so in reality error handling would become a hard-to-parse mess. For refcounted nodes, we can replicate the "pass ownership back to program on failure" logic with this patch's semantics, albeit in an ugly way: struct node n = bpf_obj_new(/* ... /), m; bpf_spin_lock(&glock); m = bpf_refcount_acquire(n); if (bpf_rbtree_add(&some_tree, &n->node, /* ... /)) { / Do something with m / } bpf_spin_unlock(&glock); bpf_obj_drop(m); bpf_refcount_acquire is used to simulate "return owning ref on failure". This should be an uncommon occurrence, though. Addition of two verifier-fixup'd args to collection inserts =========================================================== The actual bpf_obj_drop kfunc is bpf_obj_drop_impl(void , struct btf_struct_meta ), with bpf_obj_drop macro populating the second arg with 0 and the verifier later filling in the arg during insn fixup. Because bpf_rbtree_add and bpf_list_push_{front,back} now might do bpf_obj_drop, these kfuncs need a btf_struct_meta parameter that can be passed to bpf_obj_drop_impl. Similarly, because the 'node' param to those insert functions is the bpf_{list,rb}_node within the node type, and bpf_obj_drop expects a pointer to the beginning of the node, the insert functions need to be able to find the beginning of the node struct. A second verifier-populated param is necessary: the offset of {list,rb}_node within the node type. These two new params allow the insert kfuncs to correctly call __bpf_obj_drop_impl: beginning_of_node = bpf_rb_node_ptr - offset if (already_inserted) __bpf_obj_drop_impl(beginning_of_node, btf_struct_meta->record); Similarly to other kfuncs with "hidden" verifier-populated params, the insert functions are renamed with _impl prefix and a macro is provided for common usage. For example, bpf_rbtree_add kfunc is now bpf_rbtree_add_impl and bpf_rbtree_add is now a macro which sets "hidden" args to 0. Due to the two new args BPF progs will need to be recompiled to work with the new _impl kfuncs. This patch also rewrites the "hidden argument" explanation to more directly say why the BPF program writer doesn't need to populate the arguments with anything meaningful. How does this new logic affect non-owning references? ===================================================== Currently, non-owning refs are valid until the end of the critical section in which they're created. We can make this guarantee because, if a non-owning ref exists, the referent was added to some collection. The collection will drop() its nodes when it goes away, but it can't go away while our program is accessing it, so that's not a problem. If the referent is removed from the collection in the same CS that it was added in, it can't be bpf_obj_drop'd until after CS end. Those are the only two ways to free the referent's memory and neither can happen until after the non-owning ref's lifetime ends. On first glance, having these collection insert functions potentially bpf_obj_drop their input seems like it breaks the "can't be bpf_obj_drop'd until after CS end" line of reasoning. But we care about the memory not being _freed_ until end of CS end, and a previous patch in the series modified bpf_obj_drop such that it doesn't free refcounted nodes until refcount == 0. So the statement can be more accurately rewritten as "can't be free'd until after CS end". We can prove that this rewritten statement holds for any non-owning reference produced by collection insert functions: If the input to the insert function is _not_ refcounted * We have an owning reference to the input, and can conclude it isn't in any collection * Inserting a node in a collection turns owning refs into non-owning, and since our input type isn't refcounted, there's no way to obtain additional owning refs to the same underlying memory * Because our node isn't in any collection, the insert operation cannot fail, so bpf_obj_drop will not execute * If bpf_obj_drop is guaranteed not to execute, there's no risk of memory being free'd * Otherwise, the input to the insert function is refcounted * If the insert operation fails due to the node's list_head or rb_root already being in some collection, there was some previous successful insert which passed refcount to the collection * We have an owning reference to the input, it must have been acquired via bpf_refcount_acquire, which bumped the refcount * refcount must be >= 2 since there's a valid owning reference and the node is already in a collection * Insert triggering bpf_obj_drop will decr refcount to >= 1, never resulting in a free So although we may do bpf_obj_drop during the critical section, this will never result in memory being free'd, and no changes to non-owning ref logic are needed in this patch. Signed-off-by: Dave Marchevsky <davemarchevsky@fb.com> Link: https://lore.kernel.org/r/20230415201811.343116-6-davemarchevsky@fb.com Signed-off-by: Alexei Starovoitov <ast@kernel.org>		2023-04-15 17:36:50 -07:00
..
preload	bpf: iterators: Split iterators.lskel.h into little- and big- endian versions	2023-01-28 12:45:15 -08:00
arraymap.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
bloom_filter.c	bpf: compute hashes in bloom filter similar to hashmap	2023-04-02 08:44:49 -07:00
bpf_cgrp_storage.c	bpf: Teach verifier that certain helpers accept NULL pointer.	2023-04-04 16:57:16 -07:00
bpf_inode_storage.c	bpf: Teach verifier that certain helpers accept NULL pointer.	2023-04-04 16:57:16 -07:00
bpf_iter.c	bpf: implement numbers iterator	2023-03-08 16:19:51 -08:00
bpf_local_storage.c	bpf: Handle NULL in bpf_local_storage_free.	2023-04-12 10:27:50 -07:00
bpf_lru_list.c	bpf_lru_list: Read double-checked variable once without lock	2021-02-10 15:54:26 -08:00
bpf_lru_list.h	printk: stop including cache.h from printk.h	2022-05-13 07:20:07 -07:00
bpf_lsm.c	bpf: Fix the kernel crash caused by bpf_setsockopt().	2023-01-26 23:26:40 -08:00
bpf_struct_ops_types.h	bpf: Add dummy BPF STRUCT_OPS for test purpose	2021-11-01 14:10:00 -07:00
bpf_struct_ops.c	bpf: Check IS_ERR for the bpf_map_get() return value	2023-03-24 12:40:47 -07:00
bpf_task_storage.c	bpf: Teach verifier that certain helpers accept NULL pointer.	2023-04-04 16:57:16 -07:00
btf.c	bpf: Introduce opaque bpf_refcount struct and add btf_record plumbing	2023-04-15 17:36:49 -07:00
cgroup_iter.c	bpf: Pin the start cgroup in cgroup_iter_seq_init()	2022-11-21 17:40:42 +01:00
cgroup.c	bpf: allow ctx writes using BPF_ST_MEM instruction	2023-03-03 21:41:46 -08:00
core.c	bpf: Support 64-bit pointers to kfuncs	2023-04-13 21:36:41 -07:00
cpumap.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
cpumask.c	bpf: Treat KF_RELEASE kfuncs as KF_TRUSTED_ARGS	2023-03-25 16:56:22 -07:00
devmap.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
disasm.c	bpf: Relicense disassembler as GPL-2.0-only OR BSD-2-Clause	2021-09-02 14:49:23 +02:00
disasm.h	bpf: Relicense disassembler as GPL-2.0-only OR BSD-2-Clause	2021-09-02 14:49:23 +02:00
dispatcher.c	bpf: Synchronize dispatcher update with bpf_dispatcher_xdp_func	2022-12-14 12:02:14 -08:00
hashtab.c	bpf: optimize hashmap lookups when key_size is divisible by 4	2023-04-01 15:08:19 -07:00
helpers.c	bpf: Migrate bpf_rbtree_add and bpf_list_push_{front,back} to possibly fail	2023-04-15 17:36:50 -07:00
inode.c	fs: port inode_init_owner() to mnt_idmap	2023-01-19 09:24:28 +01:00
Kconfig	rcu: Make the TASKS_RCU Kconfig option be selected	2022-04-20 16:52:58 -07:00
link_iter.c	bpf: Add bpf_link iterator	2022-05-10 11:20:45 -07:00
local_storage.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
log.c	bpf: Relax log_buf NULL conditions when log_level>0 is requested	2023-04-11 18:05:44 +02:00
lpm_trie.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
Makefile	bpf: Split off basic BPF verifier log into separate file	2023-04-11 18:05:42 +02:00
map_in_map.c	bpf: Remove btf_field_offs, use btf_record's fields instead	2023-04-15 17:36:49 -07:00
map_in_map.h
map_iter.c	bpf: Introduce MEM_RDONLY flag	2021-12-18 13:27:41 -08:00
memalloc.c	bpf: Add a few bpf mem allocator functions	2023-03-25 19:52:51 -07:00
mmap_unlock_work.h	bpf: Introduce helper bpf_find_vma	2021-11-07 11:54:51 -08:00
net_namespace.c	net: Add includes masked by netdevice.h including uapi/bpf.h	2021-12-29 20:03:05 -08:00
offload.c	bpf: offload map memory usage	2023-03-07 09:33:43 -08:00
percpu_freelist.c	bpf: Initialize same number of free nodes for each pcpu_freelist	2022-11-11 12:05:14 -08:00
percpu_freelist.h
prog_iter.c
queue_stack_maps.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
reuseport_array.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
ringbuf.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
stackmap.c	bpf: return long from bpf_map_ops funcs	2023-03-22 15:11:30 -07:00
syscall.c	bpf: Introduce opaque bpf_refcount struct and add btf_record plumbing	2023-04-15 17:36:49 -07:00
sysfs_btf.c	bpf: Load and verify kernel module BTFs	2020-11-10 15:25:53 -08:00
task_iter.c	bpf: keep a reference to the mm, in case the task is dead.	2022-12-28 14:11:48 -08:00
tnum.c	bpf, tnums: Provably sound, faster, and more precise algorithm for tnum_mul	2021-06-01 13:34:15 +02:00
trampoline.c	bpf: Fix attaching fentry/fexit/fmod_ret/lsm to modules	2023-03-15 18:38:21 -07:00
verifier.c	bpf: Migrate bpf_rbtree_add and bpf_list_push_{front,back} to possibly fail	2023-04-15 17:36:50 -07:00